Optical skin detection for face unlock
Patent Information
- Application Number
- CN202280015245.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-02-18
- Filing Date
- 2022-02-17
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2042-02-17
AI Technical Summary
尽管可以检测到使用合法面部照片和视频的简单呈现攻击,但仍然缺少可靠检测使用3D面罩的呈现攻击的方法
Smart Images

Figure CN116897349B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method for facial recognition, a mobile device, and various uses of the method. The devices, methods, and uses according to the invention can be specifically used in various fields, such as: daily life, security technology, gaming, transportation technology, production technology, photography (e.g., digital or video photography for artistic, documentary, or technical purposes), security technology, information technology, agriculture, crop protection, maintenance, cosmetics, medical technology, or science. However, other applications are also possible. Background Technology
[0002] In today's digital world, secure access to information technology is a fundamental requirement for any state-of-the-art system. Standard concepts such as passphrases or PIN codes have been expanded, combined, and even replaced by biometric methods such as fingerprints or facial recognition. While passphrases can provide a high level of security if carefully chosen and of sufficient length, this step requires care and the ability to remember several potentially long phrases, depending on the IT environment. Furthermore, there is never a guarantee that the person providing the passphrase is authorized by the passphrase or digital device owner. In contrast, biometric features such as facial or fingerprint recognition are unique and individual-specific. Therefore, using features derived from these is not only more convenient but also more secure than passphrases / PIN codes, as they link personal identity with the unlocking process.
[0003] Unfortunately, similar to password theft, fingerprints and faces can also be artificially created to impersonate legitimate users. First-generation automatic facial recognition tools used digital camera images or image streams, applying 2D image processing methods to extract characteristic features, while simultaneously using machine learning techniques to generate facial templates for identification based on these features. Second-generation facial recognition algorithms use deep convolutional neural networks instead of hand-crafted image features to generate classification models.
[0004] However, both methods are vulnerable to attacks, such as those using high-quality photos that are now available for free download from the internet. Therefore, the idea of presentation attack detection (PAD) becomes significant. Early methods aimed to prevent simple attacks, such as presenting a legitimate user's photo by recording a series of images and testing for time-related features, such as subtle but natural changes in head position or blinking. These methods can again be fooled by playing pre-recorded videos of the user or by carefully animated sequences generated from publicly available photos. To exclude displays as potential targets for deception, near-infrared (NIR) cameras can be used, as displays emit photons only in the visible range of the electromagnetic spectrum. As a further countermeasure, 3D cameras have been introduced, which can clearly distinguish between a flat photograph or a tablet with a playing video and a 3D face. Nevertheless, these systems are still susceptible to attacks using high-quality masks, such as those created through 3D printing, careful 3D arrangement of 2D photographs, or handcrafted silicone or latex masks. Since humans typically wear masks, liveness detection based on small movements also fails. However, these types of masks may be rejected by the PAD system, which is capable of classifying human skin from other materials, as described in, for example, European Patent Application No. 20159984.2 filed on February 28, 2020 and European Patent Application No. 20154961.5190679 filed on January 31, 2020, the entire contents of which are incorporated herein by reference.
[0005] Another issue may stem from differences in the optical properties of skin in terms of color. Reliable identification and PAD technology require complete ignorance of these differences.
[0006] Beyond security considerations, the speed of the unlocking process and the required computing power are also crucial for providing an acceptable user experience. After successfully unlocking the device, high-speed facial recognition can be used to perform multiple tasks, such as checking if the user is still in front of the display, or initiating further security applications, such as banking apps, by performing an instant check on the person in front of the display. Again, speed and computing resources have a significant impact on the user experience.
[0007] Current 3D algorithms are computationally very demanding, and presentation attack detection requires processing multiple video frames. Therefore, expensive hardware is required to provide acceptable unlocking performance. Furthermore, this results in high power consumption.
[0008] In summary, current methods for facial unlocking cannot reliably detect spoofing attacks using 3D masks, nor can they perform the task at speeds below the limits of human detection.
[0009] US 2019 / 213309 A1 describes a system and method for authenticating a user's face using a ranging sensor. The ranging sensor includes a time-of-flight sensor and a reflectivity sensor. The ranging sensor emits a signal that is reflected away by the user and received back at the ranging sensor. The received signal can be used to determine the distance between the user and the sensor, as well as the user's reflectivity value. Using either distance or reflectivity, a processor can activate the facial recognition process in response to the distance and reflectivity.
[0010] The problem solved by this invention
[0011] Therefore, the object of this invention is to provide an apparatus and method that addresses the aforementioned technical challenges of known devices and methods. While simple presentation attacks using legitimate facial photos and videos can be detected, reliable methods for detecting presentation attacks using 3D masks are still lacking. Specifically, an additional layer of security is needed to enable the unlocking of digital devices by replacing passphrases / PIN codes with biometrics derived from the face. Furthermore, a method that is completely agnostic to different skin types from different skin tones is required. Summary of the Invention
[0012] This problem is solved by the present invention, which features independent patent claims. Advantageous developments of the invention, which can be implemented individually or in combination, are presented in the dependent claims and / or the following description and detailed embodiments.
[0013] As used below, the terms “have,” “contain,” or “include,” or any of their grammatical variations, are used in a non-exclusive manner. Therefore, these terms can refer either to a situation where no other features exist in the entity described in the context besides those introduced by these terms, or to a situation where one or more other features exist. For example, the statements “A has B,” “A contains B,” and “A includes B” can all refer to a situation where no other elements exist in A besides B (i.e., A consists entirely of B), or they can refer to a situation where entity A contains one or more other elements besides B, such as element C, elements C and D, or even other elements.
[0014] Furthermore, it should be noted that the terms "at least one," "one or more," or similar expressions indicating that a feature or element may appear once or more are generally used only once when describing the respective feature or element. In the following text, in most cases, the phrase "at least one" or "one or more" will not be repeated when referring to individual features or elements, even though the individual features or elements may appear once or more.
[0015] Furthermore, as used herein, the terms “preferredly,” “more preferably,” “particularly,” “more specifically,” “especially,” “even more particularly,” or similar terms are used in combination with optional features without limiting the possibility of alternatives. Therefore, the features introduced by these terms are optional features and are not intended to limit the scope of the claims in any way. As those skilled in the art will recognize, the invention can be practiced by using alternative features. Similarly, features introduced by phrases such as “in embodiments of the invention” are intended to be optional features and do not limit any alternative embodiments of the invention, the scope of the invention, or the possibility of combining features introduced in this way with other optional or non-optional features of the invention.
[0016] In a first aspect of the invention, a method for facial authentication is disclosed. The face to be authenticated may specifically be a human face. The term "facial authentication" as used herein is a broad term and should be given its common and conventional meaning by those skilled in the art, and is not limited to a specific or customized meaning. The term may specifically refer to, but is not limited to, verifying that an identified object or a portion of an identified object is a face. Specifically, authentication may include distinguishing a genuine human face from attack material generated to mimic a human face. Authentication may include verifying the identity of the corresponding user and / or assigning an identity to the user. Authentication may include generating and / or providing identity information, for example to other devices, such as to at least one authorized device, for authorizing access to mobile devices, machines, vehicles, buildings, etc. The identity information can be proven through authentication. For example, the identity information may be and / or may include at least one identity token. In the case of successful authentication, the identified object or a portion of the identified object is verified as a genuine face and / or the identity of the object, particularly the user, is verified.
[0017] The method includes the following steps: a) at least one face detection step, wherein the face detection step includes determining at least one first image by using at least one camera, wherein the first image includes at least one two-dimensional image of a scene suspected of including a face, wherein the face detection step includes detecting a face in the first image by using at least one processing unit to identify at least one predefined or predetermined geometric feature for the face in the first image. b) At least one skin detection step, wherein the skin detection step includes projecting at least one illumination pattern comprising a plurality of illumination features onto a scene using at least one illumination unit and determining at least one second image using at least one camera, wherein the second image includes a plurality of reflection features generated by the scene in response to illumination of the illumination features, wherein each reflection feature includes at least one beam profile, wherein the skin detection step includes: using a processing unit to determine first beam profile information of at least one reflection feature by analyzing the beam profile of at least one reflection feature among the reflection features located in an image region of the second image, the image region of the second image corresponding to an image region of the first image including identified geometric features, and determining at least one material property of the reflection feature from the first beam profile information, wherein if the material property corresponds to at least one property characteristic of skin, the detected facial feature is characterized as skin; c) At least one 3D detection step, wherein the 3D detection step includes: using a processing unit to determine second beam profile information of at least four reflection features by analyzing beam profiles of at least four reflection features located in an image region of a second image corresponding to an image region of a first image including the identified geometric features, and determining at least one depth level from the second beam profile information of the reflection features, wherein if the depth level deviates from the depth level of a planar object, the detected face is characterized as a 3D object. d) At least one authentication step, wherein the authentication step includes: authenticating the detected face by using at least one authentication unit if the face detected in step b) is characterized as skin and the face detected in step c) is characterized as a 3D object.
[0018] The method steps can be executed in a given order or in a different order. Furthermore, there may be one or more additional method steps not listed. Moreover, one, more, or even all of the method steps may be executed repeatedly.
[0019] The term "camera" as used herein is a broad term and should be given its common and conventional meaning by those skilled in the art, and is not limited to any particular or custom meaning. Specifically, the term may refer to, but is not limited to, a device having at least one imaging element configured to record or capture spatially resolved one-dimensional, two-dimensional, or even three-dimensional optical data or information. A camera may be a digital camera. As an example, a camera may include at least one camera chip, such as at least one CCD chip and / or at least one CMOS chip configured to record images. A camera may be or may include at least one near-infrared camera.
[0020] As used herein, but not limited to, the term "image" can specifically refer to data recorded using a camera, such as multiple electronic readings from an imaging device, like pixels from a camera chip. In addition to at least one camera chip or imaging chip, a camera may also include additional elements, such as one or more optical elements, such as one or more lenses. As an example, a camera may be a fixed-focus camera with at least one lens that is fixedly adjusted relative to the camera. However, alternatively, a camera may also include one or more variable lenses that can be adjusted automatically or manually.
[0021] The camera can be a camera on a mobile device, such as a laptop, tablet, or specifically a mobile phone, such as a smartphone. Therefore, specifically, the camera can be part of a mobile device that, in addition to at least one camera, includes one or more data processing devices, such as one or more data processors. However, other cameras are also possible. As used herein, the term "mobile device" is a broad term and will be given its common and conventional meaning to those skilled in the art, and is not limited to a particular or customary meaning. The term can specifically refer to, but is not limited to, mobile electronic devices, and more specifically, mobile communication devices, such as cellular phones or smartphones. Alternatively or additionally, a mobile device can also refer to a tablet computer or another type of portable computer.
[0022] Specifically, a camera may be or may include at least one optical sensor having at least one photosensitive region. As used herein, "optical sensor" generally refers to a photosensitive device for detecting a light beam, such as for detecting illumination and / or light spots generated by at least one light beam. As further used herein, "photosensitive region" generally refers to an area of the optical sensor that can be illuminated from the outside by at least one light beam, generating at least one sensor signal in response to such illumination. The photosensitive region may specifically be located on the surface of each optical sensor. However, other embodiments are also possible. A camera may include multiple optical sensors, each having a photosensitive region. As used herein, the term "optical sensor each having at least one photosensitive region" refers to a configuration of multiple individual optical sensors each having a photosensitive region and a configuration of a combined optical sensor having multiple photosensitive regions. Furthermore, the term "optical sensor" refers to a photosensitive device configured to generate an output signal. In the case where the camera includes multiple optical sensors, each optical sensor may be implemented such that a photosensitive region is precisely present in the respective optical sensor, for example by providing a precisely illuminated photosensitive region, in response to which a precisely uniform sensor signal is created for the entire optical sensor. Thus, each optical sensor may be a single-region optical sensor. However, the use of a single-area optical sensor makes camera setup particularly simple and efficient. Therefore, as an example, commercially available light sensors, such as commercially available silicon photodiodes, can be used in this setup, each with a precisely defined photosensitive area. However, other embodiments are also feasible.
[0023] Specifically, the optical sensor may be or may include at least one photodetector, preferably an inorganic photodetector, more preferably an inorganic semiconductor photodetector, and most preferably a silicon photodetector. Specifically, the optical sensor may be sensitive in the infrared spectral range. The optical sensor may include at least one sensor element comprising a matrix of pixels. All pixels of the matrix or at least one group of optical sensors in the matrix may specifically be identical. Specifically, the same group of pixels in the matrix may be provided for different spectral ranges, or all pixels may be identical in terms of spectral sensitivity. Furthermore, the size of the pixels and / or their electronic or photoelectric properties may be identical. Specifically, the optical sensor may be or may include at least one array of inorganic photodiodes sensitive in the infrared spectral range, preferably in the range of 700 nm to 3.0 micrometers. Specifically, the optical sensor may be partially sensitive in the near-infrared region suitable for silicon photodiodes, specifically in the range of 700 nm to 1100 nm. The infrared optical sensor that can be used in the optical sensor may be a commercially available infrared optical sensor, such as trinamiX from Ludwigshafen, Rhineland-sur-Holland, D-67056, Germany.TM GmbH with Hertzstueck TM Commercially available infrared optical sensors. Therefore, as an example, the optical sensor may include at least one intrinsic photovoltaic optical sensor, more preferably at least one semiconductor photodiode selected from the group consisting of: Ge photodiodes, InGaAs photodiodes, extended InGaAs photodiodes, InAs photodiodes, InSb photodiodes, and HgCdTe photodiodes. Alternatively, the optical sensor may include at least one external photovoltaic optical sensor, more preferably at least one semiconductor photodiode selected from the group consisting of: Ge:Au photodiodes, Ge:Hg photodiodes, Ge:Cu photodiodes, Ge:Zn photodiodes, Si:Ga photodiodes, and Si:As photodiodes. Alternatively, the optical sensor may include at least one photoconductivity sensor, such as a PbS or PbSe sensor, or a radiative thermal meter, preferably selected from the group consisting of: V0 radiative thermal meters and amorphous Si radiative thermal meters.
[0024] Optical sensors can be sensitive in one or more of the ultraviolet, visible, or infrared spectral ranges. Specifically, optical sensors can be sensitive in the visible spectral range of 500 nm to 780 nm, most preferably in the range of 650 nm to 750 nm, or in the range of 690 nm to 700 nm. Specifically, optical sensors can be sensitive in the near-infrared region. Specifically, optical sensors can be partially sensitive in the near-infrared region suitable for silicon photodiodes, specifically in the range of 700 nm to 1000 nm. Specifically, optical sensors can be sensitive in the infrared spectral range, specifically in the range of 780 nm to 3.0 micrometers. For example, each optical sensor can independently be or may include at least one element selected from the group consisting of: photodiodes, photoelectric elements, photoconductors, phototransistors, or any combination thereof. For example, optical sensors can be or may include at least one element selected from the group consisting of: CCD sensor elements, CMOS sensor elements, photodiodes, photoelectric cells, photoconductors, phototransistors, or any combination thereof. Any other type of photosensitive element can be used. Photosensitive elements can generally be made wholly or partially of inorganic materials and / or can be made wholly or partially of organic materials. Most commonly, one or more photodiodes can be used, such as commercially available photodiodes, such as inorganic semiconductor photodiodes.
[0025] An optical sensor may include at least one sensor element comprising a matrix of pixels. Therefore, as an example, an optical sensor may be part of or constitute a pixelated optical device. For instance, an optical sensor may be and / or may include at least one CCD and / or CMOS device. As an example, an optical sensor may be part of or constitute at least one CCD and / or CMOS device having a matrix of pixels, each pixel forming a photosensitive area.
[0026] As used herein, the term "sensor element" generally refers to a device or combination of devices configured to sense at least one parameter. Here, the parameter may specifically be an optical parameter, and the sensor element may specifically be an optical sensor element. Sensor elements can be formed as a single, integral device or as a combination of several devices. Sensor elements include optical sensor matrices. Sensor elements may include at least one CMOS sensor. The matrix may consist of individual pixels, such as individual optical sensors. Thus, a matrix of inorganic photodiodes can be formed. However, alternatively, one or more commercially available matrices, such as CCD detectors (e.g., CCD detector chips) and / or CMOS detectors (e.g., CMOS detector chips), can be used. Therefore, in general, sensor elements can be and / or may include at least one CCD and / or CMOS device and / or optical sensor that can form a sensor array or be part of a sensor array, such as the matrix described above. Thus, as an example, a sensor element may include a pixel array, such as a rectangular array with m rows and n columns, where m and n are independently positive integers. Preferably, more than one column and more than one row are given, i.e., n>1, m>1. Therefore, as an example, n can be 2 to 16 or higher, and m can be 2 to 16 or higher. Preferably, the ratio of the number of rows to the number of columns is close to 1. As an example, n and m can be chosen such that 0.3 ≤ m / n ≤ 3, for example by choosing m / n = 1:1, 4:3, 16:9, or similar ratios. As an example, the array can be a square array with the same number of rows and columns, for example by choosing m=2, n=2 or m=3, n=3, etc.
[0027] This matrix can be composed of individual pixels, such as individual optical sensors. Therefore, it can form a matrix of inorganic photodiodes. Alternatively, however, commercially available matrices can be used, such as one or more of a CCD detector (e.g., a CCD detector chip) and / or a CMOS detector (e.g., a CMOS detector chip). Therefore, typically, the optical sensor can be and / or can include at least one CCD and / or CMOS device and / or camera optical sensor forming a sensor array, or can be part of a sensor array, such as the matrix described above.
[0028] The matrix can specifically be a rectangular matrix having at least one row, preferably multiple rows, and multiple columns. As an example, the rows and columns can be substantially perpendicularly oriented. As used herein, the term "substantially perpendicular" refers to a perpendicular orientation with a tolerance of, for example, ±20° or less, preferably ±10° or less, and more preferably ±5° or less. Similarly, the term "substantially parallel" refers to a parallel orientation with a tolerance of, for example, ±20° or less, preferably ±10° or less, and more preferably ±5° or less. Therefore, as an example, tolerances less than 20°, specifically less than 10°, or even less than 5° can be acceptable. To provide a wider field of view, the matrix can specifically have at least 10 rows, preferably at least 500 rows, and more preferably at least 1000 rows. Similarly, the matrix can have at least 10 columns, preferably at least 500 columns, and more preferably at least 1000 columns. The matrix may include at least 50 optical sensors, preferably at least 100,000 optical sensors, and more preferably at least 5,000,000 optical sensors. The matrix may include multiple pixels in the megapixel range. However, other embodiments are also possible. Therefore, in a configuration where axial rotational symmetry is desired, a circular or concentric arrangement of the matrix's optical sensors (also referred to as pixels) may be preferred.
[0029] Therefore, as an example, the sensor element may be part of or constitute a pixelated optical device. For example, the sensor element may be and / or may include at least one CCD and / or CMOS device. As an example, the sensor element may be part of or constitute at least one CCD and / or CMOS device having a pixel matrix, with each pixel forming a photosensitive area. The sensor element may employ a rolling shutter or global shutter method to read out the matrix of the optical sensor.
[0030] The camera may also include at least one transfer device. The camera may also include one or more additional elements, such as one or more additional optical elements. The camera may include at least one optical element selected from the group consisting of: a transfer device, such as at least one lens and / or at least one lens system, or at least one diffractive optical element. The term "transfer device," also referred to as "transfer system," generally refers to one or more optical elements adapted to modify a light beam, for example, by modifying one or more of the beam parameters, beam width, or beam direction of the beam. The transfer device may be adapted to guide the light beam onto an optical sensor. Specifically, the transfer device may include one or more of the following: at least one lens, such as at least one lens selected from the group consisting of: at least one focusing lens, at least one aspherical lens, at least one spherical lens, at least one Fresnel lens; at least one diffractive optical element; at least one concave mirror; at least one beam deflection element, preferably at least one mirror; at least one beam splitter, preferably at least one of a beam splitter cube or beam splitter mirror; at least one multi-lens system. The transfer device may have a focal length. As used herein, the term "focal length" of a transfer device refers to the distance from which an incident collimated ray of light that may impact the transfer device enters the "focal point," which can also be interpreted as the "focus point." Therefore, the focal length constitutes a measure of the transfer device's ability to converge an incident beam of light. Consequently, a transfer device may include one or more imaging elements that can have the effect of a converging lens. For example, a transfer device may have one or more lenses, particularly one or more refracting lenses, and / or one or more convex mirrors. In this example, the focal length may be defined as the distance from the center of a thin refracting lens to the principal focal point of the thin lens. For converging thin refracting lenses, such as convex or biconvex thin lenses, the focal length can be considered positive and can provide the distance from which a collimated beam of light illuminates the thin lens when the transfer device can be focused into a single point of light. Additionally, the transfer device may include at least one wavelength selection element, such as at least one filter. Furthermore, the transfer device may be designed to impose a predetermined beam profile on electromagnetic radiation (e.g., at the sensor region and, particularly, at the sensor region's location). In principle, the above-described alternative embodiments of the transfer device can be implemented individually or in any desired combination.
[0031] A transmission device may have an optical axis. As used herein, the term "optical axis of a transmission device" generally refers to the mirror symmetry axis or rotational symmetry axis of a lens or lens system. As an example, a transmission system may include at least one beam path, wherein elements of the transmission system within the beam path are positioned in a rotationally symmetric manner relative to the optical axis. However, one or more optical elements located within the beam path may also be eccentric or tilted relative to the optical axis. In this case, however, the optical axis may be defined sequentially, for example, through the center of the optical elements in the interconnected beam paths, such as through the center of the interconnected lenses, wherein, in this case, the optical sensor is not considered an optical element. An optical axis may generally represent a beam path. A camera may have a single beam path along which a beam travels from an object to an optical sensor, or it may have multiple beam paths. As an example, a single beam path may be given, or the beam path may be divided into two or more partial beam paths. In the latter case, each partial beam path may have its own optical axis. In the case of multiple optical sensors, the optical sensors may be located in the same beam path or partial beam paths. However, alternatively, the optical sensors may also be located in different partial beam paths.
[0032] The transmission device can form a coordinate system, where the longitudinal coordinate is the coordinate along the optical axis, and where d is the spatial offset from the optical axis. The coordinate system can also be a polar coordinate system, where the optical axis of the transmission device forms the z-axis, and where the distance from the z-axis and the polar angle can be used as additional coordinates. Directions parallel to or antiparallel to the z-axis can be considered longitudinal directions, and coordinates along the z-axis can be considered longitudinal coordinates. Any direction perpendicular to the z-axis can be considered transverse directions, and polar coordinates and / or polar angles can be considered transverse coordinates.
[0033] The camera is configured to determine at least one image of a scene, specifically a first image. As used herein, the term "scene" can refer to a spatial region. The scene may include a face and its surrounding environment during authentication. The first image itself may include pixels, the pixels of which are related to the pixels of a matrix of sensor elements. Therefore, when referring to "pixel," it either refers to a unit of image information generated by a single pixel of the sensor element or directly refers to a single pixel of the sensor element. The first image is at least one two-dimensional image. As used herein, the term "two-dimensional image" can generally refer to an image having information about lateral coordinates, such as dimensions of height and width. The first image may be an RGB (red, green, blue) image. The term "determine at least one first image" may refer to capturing and / or recording the first image.
[0034] The face detection step includes detecting a face in the first image by using at least one processing unit to identify at least one predefined or predetermined geometric feature of the face in the first image. Specifically, the face detection step includes detecting a face in the first image by using at least one processing unit to identify at least one predefined or predetermined geometric feature as a facial feature in the first image.
[0035] As used further herein, the term "processing unit" generally refers to any data processing device suitable for performing specified operations, for example, by using at least one processor and / or at least one application-specific integrated circuit (ASIC). Thus, by way of example, at least one processing unit may include software code stored thereon, which includes a plurality of computer commands. The processing unit may provide one or more hardware elements for performing one or more specified operations and / or may provide one or more processors on which software for performing one or more specified operations runs. Operations including evaluating images may be performed by at least one processing unit. Thus, by way of example, one or more instructions may be implemented in software and / or hardware. Thus, by way of example, the processing unit may include one or more programmable devices, such as one or more computers, application-specific integrated circuits (ASICs), digital signal processors (DSPs), or field-programmable gate arrays (FPGAs), configured to perform the aforementioned evaluation. However, additionally or alternatively, the processing unit may also be implemented entirely or partially in hardware. The processing unit and the camera may be fully or partially integrated into a single device. Thus, typically, the processing unit may also form part of the camera. Alternatively, the processing unit and the camera may be embodied entirely or partially as separate devices.
[0036] The processing unit may be or may include one or more integrated circuits, such as one or more application-specific integrated circuits (ASICs), and / or one or more data processing devices, such as one or more computers, preferably one or more microcomputers and / or microcontrollers, field-programmable arrays, or digital signal processors. Additional components may be included, such as one or more preprocessing devices and / or data acquisition devices, such as one or more devices for receiving and / or preprocessing sensor signals, such as one or more AD converters and / or one or more filters. Furthermore, the processing unit may include one or more measuring devices, such as one or more measuring devices for measuring current and / or voltage. Additionally, the processing unit may include one or more data storage devices. Furthermore, the processing unit may include one or more interfaces, such as one or more wireless interfaces and / or one or more wired interfaces.
[0037] The processing unit can be configured to display, visualize, analyze, distribute, transmit, or further process information (such as information acquired by a camera) or other similar functions. As an example, the processing unit can be connected to or integrated with at least one of the following: a display, projector, monitor, LCD, TFT, speaker, multi-channel audio system, LED pattern, or other visualization device. It can also be connected to or integrated with at least one of the following: a communication device or communication interface, connector, or port capable of sending encrypted or unencrypted information using one or more of the following: email, text message, telephone, Bluetooth, Wi-Fi, infrared, or internet interfaces, ports, or connections. It can be further connected to or integrated with at least one of the following: a processor, graphics processor, CPU, Open Multimedia Application Platform (OMAP™), integrated circuit, system-on-a-chip (e.g., products from Apple A-series or Samsung S3C2 series), microcontroller or microprocessor, one or more memory blocks (such as ROM, RAM, EEPROM, or flash memory), timing source (such as an oscillator or phase-locked loop, counter timer, real-time timer, or power-on reset generator), voltage regulator, power management circuitry, or DMA controller. The units can also be connected via a bus (such as an AMBA bus) or integrated into IoT or Industry 4.0 type networks.
[0038] The processing unit can be connected to or has additional external interfaces or ports, such as one or more of the following: serial or parallel interfaces or ports, USB, Centronics ports, FireWire, HDMI, Ethernet, Bluetooth, RFID, Wi-Fi, USART, or SPI; or analog interfaces or ports, such as one or more ADCs or DACs; or standardized interfaces or ports to other devices, such as 2D camera devices using RGB interfaces (e.g., CameraLink). The processing unit can also be connected via inter-processor interfaces or ports, FPGA-FPGA interfaces, or one or more serial or parallel interface ports. The processing unit can also be connected to one or more of the following: optical disc drives, CD-RW drives, DVD+RW drives, flash memory drives, memory cards, disk drives, hard disk drives, solid-state drives, or solid-state drives.
[0039] The processing unit may be connected to or have one or more additional external connectors, such as one or more of the following: telephone connector, RCA connector, VGA connector, dual-pole connector, USB connector, HDMI connector, 8P8C connector, BCN connector, IEC 60320 C14 connector, fiber optic connector, D-miniature connector, RF connector, coaxial connector, SCART connector, XLR connector, and / or may include at least one suitable receptacle for one or more of these connectors.
[0040] Detecting a face in a first image may include identifying at least one predefined or predetermined geometric feature of the face. The term "geometric feature of the face" as used herein is a broad term and should be given its common and conventional meaning by those skilled in the art, and is not limited to a specific or customary meaning. Specifically, the term may refer to, but is not limited to, at least one geometry-based feature that describes the shape of the face and its components, particularly one or more of the nose, eyes, mouth, or eyebrows. The processing unit may include at least one database in which the geometric feature of the face is stored, such as in a lookup table. Techniques for identifying at least one predefined or predetermined geometric feature of the face are generally known to those skilled in the art. For example, face detection may be performed as described in Masi, Lacopo et al., “Deep face recognition: A survey,” 31st SIBGRAPI Conference on Graphics, Patterns and Images (SIBGRAPI), IEEE, 2018, the entire contents of which are incorporated herein by reference.
[0041] The processing unit can be configured to perform at least one image analysis and / or image processing to identify geometric features. Image analysis and / or image processing may use at least one feature detection algorithm. Image analysis and / or image processing may include one or more of the following: filtering; selecting at least one region of interest; background correction; decomposition into color channels; decomposition into hue, saturation, and / or brightness channels; frequency decomposition; singular value decomposition; applying a speckle detector; applying an angle detector; applying a determinant of a Hessian filter; applying a region detector based on principle curvature; applying a gradient position and orientation histogram algorithm; applying a histogram of orientation gradient descriptors; applying an edge detector; applying a differential edge detector; applying a Canny edge detector; applying a Laplacian Gaussian filter; applying a difference Gaussian filter; applying the Sobel operator; applying the Laplacian operator; applying the Scharr operator; applying the Prewitt operator; applying the Roberts operator; applying the Kirsch operator; applying a high-pass filter; applying a low-pass filter; applying a Fourier transform; applying the Radon transform; applying the Hough transform; applying the wavelet transform; thresholding; creating a binary image. The region of interest can be determined manually by the user or automatically, for example, by identifying features within the first image.
[0042] Specifically, after the face detection step, a skin detection step can be performed, including projecting at least one lighting pattern comprising multiple lighting features onto the scene using at least one lighting unit. However, embodiments in which the skin detection step is performed before the face detection step are feasible.
[0043] As used herein, the term "lighting unit," also meaning a light source, can generally refer to at least one arbitrary device configured to generate at least one lighting pattern. The lighting unit can be configured to provide a lighting pattern for illuminating a scene. The lighting unit can be adapted to directly or indirectly illuminate the scene, wherein the lighting pattern is affected (remitted) by the surface of the scene, particularly by reflection or scattering, and thereby at least partially directed towards a camera. The lighting unit can be configured to illuminate the scene, for example, by directing a light beam toward the scene and reflecting that beam. The lighting unit can be configured to generate a light beam for illuminating the scene.
[0044] The illumination unit may include at least one light source. The illumination unit may include multiple light sources. The illumination unit may include artificial lighting sources, particularly at least one laser source and / or at least one incandescent lamp and / or at least one semiconductor light source, such as at least one light-emitting diode, particularly an organic light-emitting diode and / or an inorganic light-emitting diode. As an example, the light emitted by the illumination unit may have a wavelength of 300 to 1100 nm, particularly 500 to 1100 nm. Alternatively or additionally, light in the infrared spectral range may be used, for example, light in the range of 780 nm to 3.0 μm. Specifically, light in the near-infrared region, where silicon photodiodes are particularly suitable, may be used, particularly light in the range of 700 nm to 1100 nm.
[0045] The illumination unit can be configured to generate at least one illumination pattern in the infrared region. The illumination feature can have a wavelength in the near-infrared (NIR) range. The illumination feature can have a wavelength of approximately 940 nm. At this wavelength, melanin absorption is depleted, so the dark and bright complexes reflect almost the same light. However, other wavelengths in the NIR region are also possible, such as one or more of 805 nm, 830 nm, 835 nm, 850 nm, 905 nm, or 980 nm. Furthermore, using light in the near-infrared region makes the light undetectable to the human eye or only weakly detectable, but still detectable by silicon sensors, particularly standard silicon sensors.
[0046] The illumination unit can be configured to emit light of a single wavelength. In other embodiments, the illumination unit can be configured to emit light with multiple wavelengths, thereby allowing for additional measurements in other wavelength channels.
[0047] As used herein, the term "ray" generally refers to a line perpendicular to the wavefront of light and pointing in the direction of energy flow. As used herein, the term "bundle" generally refers to a collection of rays. Hereinafter, the terms "ray" and "bundle" will be used synonymously. As further used herein, the term "beam" generally refers to the amount of light, specifically the amount of light traveling substantially in the same direction, including the possibility that the beam has an extension angle or a widening angle. A beam can have spatial extension. Specifically, a beam can have a non-Gaussian beam profile. The beam profile can be selected from the group consisting of: trapezoidal beam profile; triangular beam profile; and conical beam profile. A trapezoidal beam profile can have a plateau region and at least one edge region. Specifically, a beam can be a Gaussian beam or a linear combination of Gaussian beams, as will be further detailed below. However, other embodiments are also possible.
[0048] The illumination unit may be or may include at least one multi-beam source. For example, the illumination unit may include at least one laser source and one or more diffractive optical elements (DOEs). Specifically, the illumination unit may include at least one laser and / or laser source. Various types of lasers may be employed, such as semiconductor lasers, dual heterostructure lasers, external cavity lasers, split-confined heterostructure lasers, quantum cascade lasers, distributed Bragg reflector lasers, polaron lasers, hybrid silicon lasers, extended cavity diode lasers, quantum dot lasers, bulk Bragg grating lasers, indium arsenide lasers, transistor lasers, diode-pumped lasers, distributed feedback lasers, quantum well lasers, interband cascade lasers, gallium arsenide lasers, semiconductor ring lasers, extended cavity diode lasers, or vertical cavity surface-emitting lasers. Alternatively or additionally, non-laser source materials, such as LEDs and / or bulbs, may be used. The illumination unit may include one or more diffractive optical elements (DOEs) suitable for generating illumination patterns. For example, the illumination unit may be adapted to generate and / or project point clouds, and may include one or more of the following: at least one digital light processing projector, at least one LCoS projector, at least one spatial light modulator; at least one diffractive optical element; at least one light-emitting diode array; at least one laser source array. Using at least one laser source as the illumination unit is particularly preferred due to its typically defined beam profile and other properties of operability. The illumination unit may be integrated into the camera housing or may be detached from the camera.
[0049] Furthermore, the illumination unit can be configured to emit modulated or unmodulated light. When using multiple illumination units, the different units can have different modulation frequencies, which can later be used to distinguish the light beams.
[0050] One or more light beams generated by the illumination unit can typically propagate parallel to or oblique to the optical axis, for example, at an angle to the optical axis. The illumination unit can be configured such that one or more light beams propagate from the illumination unit toward the scene along the optical axis of the illumination unit and / or the camera. For this purpose, the illumination unit and / or the camera may include at least one reflective element, preferably at least one prism, for deflecting the illumination beam onto the optical axis. As an example, one or more light beams, such as a laser beam, may have an angle of less than 10°, preferably less than 5°, or even less than 2° with respect to the optical axis. However, other embodiments are also possible. Furthermore, one or more light beams may be on or off the optical axis. As an example, one or more light beams may be parallel to the optical axis and at a distance of less than 10 mm from the optical axis, preferably less than 5 mm, or even less than 1 mm, or may even coincide with the optical axis.
[0051] As used herein, the term "at least one lighting pattern" refers to at least one arbitrary pattern that includes at least one lighting feature suitable for at least a portion of a lighting scene. As used herein, the term "lighting feature" refers to at least one feature of a pattern that is at least partially extended. A lighting pattern may include a single lighting feature. A lighting pattern may include multiple lighting features. A lighting pattern may be selected from the group consisting of: at least one dot pattern; at least one line pattern; at least one stripe pattern; at least one checkerboard pattern; at least one pattern including an arrangement of periodic or aperiodic features. A lighting pattern may include regular and / or constant and / or periodic patterns, such as triangular patterns, rectangular patterns, hexagonal patterns, or patterns including additional convex patchwork. A lighting pattern may present at least one lighting feature selected from the group consisting of: at least one dot; at least one row; at least two lines, such as parallel or intersecting lines; at least one dot and one line; at least one arrangement of periodic or aperiodic features; at least one feature of arbitrary shape. The illumination pattern may include at least one pattern selected from the group consisting of: at least one point pattern, particularly a pseudo-random point pattern; a random point pattern or a quasi-random pattern; at least one Sobol pattern; at least one quasi-periodic pattern; at least one pattern including at least one known feature; at least one regular pattern; at least one triangular pattern; at least one hexagonal pattern; at least one rectangular pattern, including at least one pattern with convex uniform tiling; at least one line pattern including at least one line; at least one line pattern including at least two lines (e.g., parallel or intersecting lines). For example, the illumination unit may be adapted to generate and / or project a point cloud. The illumination unit may include at least one light projector adapted to generate a point cloud such that the illumination pattern may include multiple point patterns. The illumination pattern may include a periodic laser spot grid. The illumination unit may include at least one mask adapted to generate the illumination pattern from at least one beam of light generated by the illumination unit.
[0052] The distance between two features of the illumination pattern and / or the area of at least one illumination feature may depend on the circle of confusion in the image. As described above, the illumination unit may include at least one light source configured to generate at least one illumination pattern. Specifically, the illumination unit includes at least one laser source and / or at least one laser diode designated for generating laser radiation. The illumination unit may include at least one diffractive optical element (DOE). The illumination unit may include at least one point projector, such as at least one laser source and DOE, adapted to project at least one periodic point pattern. As further used herein, the term "projecting at least one illumination pattern" may refer to providing at least one illumination pattern for illuminating at least one scene.
[0053] The skin detection step includes determining at least one second image, also referred to as a reflection image, using a camera. The method may include determining multiple second images. The reflection features of the multiple second images may be used for skin detection in step b) and / or for 3D detection in step c).
[0054] The second image includes multiple reflection features generated by the scene in response to illumination features. As used herein, the term "reflection feature" can refer to a feature in an image plane generated by the scene in response to illumination (specifically having at least one illumination feature). Each reflection feature includes at least one beam profile, also referred to as a reflected beam profile. As used herein, the term "beam profile" of a reflection feature can generally refer to at least one intensity distribution of the reflection feature, such as the intensity distribution of a light spot on an optical sensor, as a function of pixels. The beam profile can be selected from the group consisting of: trapezoidal beam profiles; triangular beam profiles; conical beam profiles; and linear combinations of Gaussian beam profiles.
[0055] Evaluation of the second image may include identifying reflection features of the second image. The processing unit may be configured to perform at least one image analysis and / or image processing to identify reflection features. Image analysis and / or image processing may use at least one feature detection algorithm. Image analysis and / or image processing may include one or more of the following: filtering; selecting at least one region of interest; forming a difference image between an image created by a sensor signal and at least one offset; inverting a sensor signal by inverting an image created by a sensor signal; forming a difference image between images created by a sensor signal at different times; background correction; decomposition into color channels; decomposition into hue, saturation, and luminance channels; frequency decomposition; singular value decomposition; applying a speckle detector; applying a corner detector; applying the determinant of a Hessian filter; applying a region detector based on principle curvature; applying a maximum stable extremum region detector; applying a generalized Hough transform; applying a ridge detector; applying an affine invariant feature detector; applying an affine adaptive interest point operator; applying a Harris affine region detector; applying Hessia... n-Affine Region Detector; Applying Scale-Invariant Feature Transform; Applying Scale-Space Extremum Detector; Applying Local Feature Detector; Applying Accelerated Robust Feature Algorithm; Applying Gradient Location and Orientation Histogram Algorithm; Applying Histogram with Orientation Gradient Descriptor; Applying Deriche Edge Detector; Applying Differential Edge Detector; Applying Spatiotemporal Interest Point Detector; Applying Moravec Corner Detector; Applying Canny Edge Detector; Applying Gaussian Laplacian Filter; Applying Gaussian Difference Filter; Applying Sobel Operator; Applying Laplacian Operator; Applying Scharr Operator; Applying Prewitt Operator; Applying Roberts Operator; Applying Kirsch Operator; Applying High-Pass Filter; Applying Low-Pass Filter; Applying Fourier Transform; Applying Radon Transform; Applying Hough Transform; Applying Wavelet Transform; Thresholding; Creating Binary Images. Regions of interest can be manually determined by the user or automatically, for example, by identifying features within an image generated by an optical sensor.
[0056] For example, the illumination unit can be configured to generate and / or project point clouds, such that multiple illuminated regions are generated on an optical sensor (e.g., a CMOS detector). Additionally, interference may exist on the optical sensor, such as interference due to speckles and / or external light and / or multiple reflections. The processing unit can be adapted to determine at least one region of interest, such as one or more pixels illuminated by a light beam, for determining the longitudinal coordinates of the corresponding reflective features, which will be described in more detail below. For example, the processing unit can be adapted to perform filtering methods, such as speckle analysis and / or edge filtering and / or object recognition methods.
[0057] The processing unit can be configured to perform at least one image correction. Image correction may include at least one background subtraction. The processing unit may be adapted to remove the effects of background light from the beam profile, for example, through imaging without further illumination.
[0058] The processing unit can be configured to determine the beam profile of a corresponding reflection feature. As used herein, the term "determine beam profile" refers to identifying at least one reflection feature provided by an optical sensor and / or selecting at least one reflection feature provided by an optical sensor and evaluating at least one intensity distribution of that reflection feature. As an example, the intensity distribution, such as a three-dimensional intensity distribution or a two-dimensional intensity distribution, can be determined using a region of an evaluation matrix, for example, along an axis or line passing through the matrix. As an example, the illumination center of the beam can be determined, for example, by determining at least one pixel with the highest illumination, and a cross-sectional axis passing through the illumination center can be selected. The intensity distribution can be an intensity distribution that is a function of coordinates along a cross-sectional axis passing through the illumination center. Other evaluation algorithms are also feasible.
[0059] The reflection feature may cover at least one pixel of the second image or may extend over at least one pixel of the second image. For example, the reflection feature may cover multiple pixels or may extend over multiple pixels. The processing unit may be configured to determine and / or select all pixels connected to and / or belonging to the reflection feature (e.g., a light spot). The processing unit may be configured to determine the intensity center by: , Among them, R coi It is the location of the intensity center, r pixel It is the pixel position, and Where j is the number of pixels connected to and / or belonging to the reflection feature, and I total It is the total intensity.
[0060] The processing unit is configured to determine first beam profile information of at least one reflection feature by analyzing the beam profile of at least one reflection feature located within an image region of a second image corresponding to an image region of a first image including the identified geometric features. The method may include identifying an image region of the second image corresponding to an image region of the first image including the identified geometric features. Specifically, the method may include matching pixels of the first and second images and selecting pixels of the second image corresponding to the image region of the first image including the identified geometric features. The method may also include additional reflection features located outside the image region of the second image.
[0061] As used herein, the term "beam profile information" can refer to any information and / or properties derived from and / or associated with the beam profile of a reflection feature. The first and second beam profile information may be the same or different. For example, the first beam profile information may be an intensity distribution, reflection profile, intensity center, or material features. For skin detection in step b), beam profile analysis can be used. Specifically, beam profile analysis classifies materials using the reflection properties of coherent light projected onto the object surface. Material classification can be performed as described in WO 2020 / 187719, EP application 20159984.2 filed February 28, 2020, and / or EP application 20154961.5 filed January 31, 2020, the entire contents of which are incorporated herein by reference. Specifically, a periodic laser spot grid, such as the hexagonal grid described in EP application 20170905.2 filed April 22, 2020, is projected, and a reflection image is recorded using a camera. Feature-based methods can be used to analyze the beam profile of each reflective feature recorded by the camera. Feature-based methods are explained below. They can be combined with machine learning methods that allow for the parameterization of skin classification models. Alternatively or in combination, convolutional neural networks can be used to classify skin using reflective images as input.
[0062] Other methods for verifying a user's face are known, for example, from US 2019 / 213309 A1. However, these methods use time-of-flight (ToF) sensors. ToF sensors are known to work by emitting light and measuring the time span until the reflected light is received. In contrast, the proposed beam profilometry analysis uses a projected illumination pattern. Using such a projection pattern is not feasible for ToF sensors. Using an illumination pattern could be advantageous, for example, considering coverage and thus allowing for consideration of different locations on the face. This could enhance the reliability and security of user facial authentication.
[0063] The skin detection step may include determining at least one material property of a reflective feature from beam profile information using a processing unit. Specifically, the processing unit is configured to identify reflective features that would be generated by irradiating biological tissue, particularly human skin, if the reflected beam profile satisfies at least one predetermined or predefined criterion. As used herein, the term "at least one predetermined or predefined criterion" refers to at least one property and / or value suitable for distinguishing biological tissue (particularly human skin) from other materials. The predetermined or predefined criterion may be or may include at least one predetermined or predefined value and / or threshold and / or threshold range relating to material properties. If the reflected beam profile satisfies at least one predetermined or predefined criterion, the reflective feature may be indicated as being generated by biological tissue. As used herein, the term "indication" refers to any indication, such as an electronic signal and / or at least one visual or auditory indication. The processing unit is configured to otherwise identify the reflective feature as non-skin. As used herein, the term "biological tissue" generally refers to biological material containing living cells. Specifically, the processing unit may be configured for skin detection. The term "identification" derived from biological tissue (particularly human skin) can refer to determining and / or verifying whether a surface to be examined or tested is or includes biological tissue (particularly human skin), and / or distinguishing biological tissue (particularly human skin) from other tissues (particularly other surfaces). The method according to the invention allows for the differentiation of human skin from one or more of inorganic tissues, metallic surfaces, plastic surfaces, foams, paper, wood, displays, screens, and fabrics. The method according to the invention also allows for the differentiation of human biological tissue from surfaces of artificial or inanimate objects.
[0064] The processing unit can be configured to determine the material property m of the surface emitting the reflection feature by evaluating the beam profile of the reflection feature. As used herein, the term "material property" refers to at least one arbitrary property of a material configured to characterize and / or identify and / or classify the material. For example, a material property can be a property selected from the group consisting of: roughness, depth of light penetration into the material, properties characterizing the material as a biological or non-biological material, reflectivity, specular reflectivity, diffuse reflectivity, surface properties, measurement of translucency, scattering, particularly backscattering behavior, etc. At least one material property can be a property selected from the group consisting of: scattering coefficient, translucency, transparency, deviation of Lambertian surface reflection, speckle, etc. As used herein, the term "determining at least one material property" can refer to assigning a material property to a corresponding reflection feature, particularly to a detected face. The processing unit can include at least one database comprising a list and / or table of predefined and / or pre-determined material properties, such as a lookup list or lookup table. The list and / or table of material properties can be determined and / or generated by performing at least one test measurement, for example, by performing a material test using a sample with known material properties. The list and / or table of material properties may be determined and / or generated by the manufacturer and / or by the user. Material properties may also be assigned to a material classifier, such as one or more of the following: material name, material group, e.g., biological or non-biological material, translucent or non-translucent material, metallic or non-metallic, skin or non-skin, fur or non-fur, carpet or non-carpet, reflective or non-reflective, specular or non-specular, foam or non-foam, hair or non-hair, roughness group, etc. The processing unit may include at least one database comprising lists and / or tables that include material properties and associated material names and / or material groups.
[0065] The reflective properties of skin can be characterized by the simultaneous occurrence of direct surface reflection (Lambertian-like) and subsurface scattering (volume scattering). Compared to the materials mentioned above, this results in a wider laser spot on the skin.
[0066] The first beam profile information can be the reflection profile. For example, not wanting to be bound by this theory, human skin can have a reflection profile, also represented as a backscattering profile, comprising the portion generated by backscattering of the surface (represented as surface reflection) and the portion generated by very diffuse reflection of light penetrating the skin, represented as the diffuse portion of backscattering. For the reflection profile of human skin, see “Lasertechnik in der Medizin: Grundlagen, Systeme, Anwendungen”, “Wirkung von Laserstrahlung auf Gewebe”, 1991, pp. 10 171-266, Jürgen Eichler, Theo Seiler, Springer Verlag, ISBN 0939-0979. The surface reflection of skin may increase with increasing wavelength towards the near-infrared direction. Furthermore, the penetration depth may increase with increasing wavelength from visible light to near-infrared. The diffuse portion of backscattering may increase with the penetration depth of light. By analyzing the backscattering distribution, these properties can be used to distinguish skin from other materials.
[0067] Specifically, the processing unit may be configured to compare the reflected beam profile with at least one predetermined and / or pre-recorded and / or predefined beam profile. The predetermined and / or pre-recorded and / or predefined beam profile may be stored in a table or lookup table and may be determined, for example, empirically, and may be stored, as an example, in at least one data storage device of the detector. For example, the predetermined and / or pre-recorded and / or predefined beam profile may be determined during the initial startup of the device performing the method according to the invention. For example, the predetermined and / or pre-recorded and / or predefined beam profile may be stored, for example, by software, specifically by an application downloaded from an application store, etc., in at least one data storage device of the processing unit or device. If the reflected beam profile is identical to the predetermined and / or pre-recorded and / or predefined beam profile, the reflection feature may be identified as being generated by biological tissue. The comparison may include overlapping the reflected beam profile and the predetermined or predefined beam profile such that their intensity centers match. The comparison may include determining the deviation between the reflected beam profile and the predetermined and / or pre-recorded and / or predefined beam profile, such as the sum of squares of point-to-point distances. The processing unit may be adapted to compare the determined deviation with at least one threshold, wherein if the determined deviation is less than and / or equal to the threshold, the surface is indicated as biological tissue and / or the detection of biological tissue is confirmed. The threshold may be stored in a table or lookup table and may be determined, for example, empirically, and may, as an example, be stored in at least one data storage device of the processing unit.
[0068] Alternatively, the first beam profile information can be determined by applying at least one image filter to the image of the region. As used further herein, the term "image" refers to a two-dimensional function f(x,y), where a brightness and / or color value is given for any x,y location in the image. This location may correspond to a discretized recording pixel. The brightness and / or color may correspond to the bit depth of an optical sensor. As used herein, the term "image filter" refers to at least one mathematical operation applied to at least one specific region of the beam profile and / or beam profile. Specifically, the image filter Ф maps the image f or the region of interest in the image to the real number Ф(f(x,y)) = ,in Representing features, especially material features. Images can be affected by noise, and so can features. Therefore, features may be random variables. These features can be normally distributed. If features are not normally distributed, they can be transformed to a normal distribution, for example, through the Box-Cox transformation.
[0069] The processing unit can be configured to determine at least one material feature by applying at least one material-related image filter Ф2 to an image. 2m As used herein, the term "material-correlated" image filter refers to an image with material-correlated output. The output of a material-correlated image filter is denoted herein as "material feature". 2m "or material-related characteristics" 2m The material features may be, or may include, at least one piece of information about at least one material property of the surface of the scene for which the reflective features have been generated.
[0070] Material-related image filters can be at least one filter selected from the group consisting of: brightness filters; point shape filters; square norm gradients; standard deviations; smoothing filters, such as Gaussian filters or median filters; contrast filters based on gray-level occurrence; energy filters based on gray-level occurrence; homogeneity filters based on gray-level occurrence; dissimilarity filters based on gray-level occurrence; law-based energy filters; threshold region filters; or linear combinations thereof; or further material-related image filters. 2otherIt is related to one or more of the following: brightness filter, spot shape filter, square norm gradient, standard deviation, smoothness filter, energy filter based on gray level occurrence, uniformity filter based on gray level occurrence, dissimilarity filter based on gray level occurrence, law-based energy filter, or threshold region filter, or a linear combination thereof |ρ Ф2other,Фm |≥0.40, where Ф m This can be one or more of the following: brightness filter, spot shape filter, square norm gradient, standard deviation, smoothing filter, energy filter based on gray level occurrence, uniformity filter based on gray level occurrence, dissimilarity filter based on gray level occurrence, law-based energy filter, or threshold region filter, or a linear combination thereof. Another material-related image filter Ф 2other It can be done through |ρ Ф2other,Фm |≥0.60, preferably through |ρ Ф2other,Фm |≥0.80 Material-related image filter Ф m One or more related ones.
[0071] The material-dependent image filter can be at least one arbitrary filter Φ that passes the hypothesis test. As used herein, the term "passes the hypothesis test" means the fact that the null hypothesis H0 is rejected and the alternative hypothesis H1 is accepted. Hypothesis testing can include testing the material-dependent nature of the image filter by applying it to a predefined dataset. This dataset can include multiple beam profile images. As used herein, the term "beam profile image" refers to... Sum of Gaussian radial basis functions ,
[0072] Each of them Gaussian radial basis functions are centered Pre-factor and index factors Let's define it. For all Gaussian functions across all images, the exponential factor is the same. For all images... central position They are all the same: , Each beam profile image in the dataset can correspond to a material classifier and a distance. The material classifier can be labeled "Material A", "Material B", etc. The above formula can be used. The beam profile image is generated by combining the following parameter table:
[0073] The values of x and y correspond to having The integer value of the pixel. The image can have a pixel size of 32x32. The dataset of beam contour images can be obtained by using the formula above. And combine it with the parameter set to generate, in order to obtain A continuous description. The value of each pixel in a 32x32 image can be obtained through... China's target , To obtain the value, insert an integer value from 0, ..., 31. For example, for pixel (6, 9), the value can be calculated. .
[0074] Then, for each image It is possible to calculate the eigenvalues corresponding to filter Φ. , , where z k It corresponds to an image from a predefined dataset. Distance value. This will produce corresponding generated feature values. The dataset. Hypothesis testing can use the null hypothesis, that the filter does not distinguish between material classifiers. The null hypothesis can be given by H0: ,in It corresponds to the eigenvalue Expected value for each material group. Index This represents a group of materials. Hypothesis testing can use alternative hypotheses that the filter can indeed distinguish at least two materials. The alternative hypothesis can be derived from H1: As used herein, the term "non-discriminating material classifier" means that the expected values of the material classifiers are the same. As used herein, the term "discriminating material classifier" means that at least two of the expected values of the material classifiers are different. As used herein, "discriminating at least two material classifiers" is synonymous with "suitable material classifier". Hypothesis testing may include at least one analysis of variance (ANOVA) on the generated eigenvalues. Specifically, hypothesis testing may include determining each The average value of the material's characteristic values, i.e., the total average value, ,for ,in, Give each of the following in a predefined dataset The number of eigenvalues of the material. Hypothesis testing may include determining the average of all N eigenvalues. Hypothesis testing may include determining the mean square sum of the following: .
[0075] Hypothesis testing may include determining the mean square sum among the following: .
[0076] Hypothesis testing may include performing an F-test: ,in, , F( ) = 1 –
[0077] p = F(
[0078] here, It is a regularized, incomplete Beta function. Among them, Euler's Beta function and This is an incomplete Beta function. The image filter passes the hypothesis test if the p-value is less than or equal to a predefined significance level. The filter passes the hypothesis test if p ≤ 0.075, preferably p ≤ 0.05, more preferably p ≤ 0.025, and most preferably p ≤ 0.01. For example, with a predefined significance level of α = 0.075, the image filter passes the hypothesis test if the p-value is less than α = 0.075. In this case, the null hypothesis H0 can be rejected and the alternative hypothesis H1 can be accepted. The image filter thus distinguishes at least two material classifiers. Therefore, the image filter passes the hypothesis test.
[0079] Below, we assume that the reflected image includes at least one reflection feature, specifically a spot image, to describe the image filter. The spot image f can be described by the function In this case, the background of image f may have been subtracted. However, other reflection features are also possible.
[0080] For example, a material-related image filter can be a brightness filter. A brightness filter can return a measurement of the brightness of a light spot as a material feature. The material feature can be determined by the following formula: , Where f is the spot image. The distance of the spot is represented by z, where z can be obtained, for example, by using depth-from-photon ratio techniques or / or by using triangulation techniques. The surface normal of the material is... This vector is given and can be obtained as the normal to the surface spanned by at least three measured light spots. It is the direction vector of the light source. The position of the light spot is determined by using defocus depth or photon depth ratio techniques and / or by using triangulation techniques, where the position of the light source is referred to as a parameter of the detector system. , is the difference vector between the position of the light spot and the position of the light source.
[0081] For example, a material-related image filter can be a filter with an output that depends on the shape of the light spot. This material-related image filter can return a value related to the material's translucency as a material feature. The material's translucency affects the shape of the light spot. The material feature can be given by the following formula: , in, These are the weights of the light spot height h, and H represents the Heavyside function, i.e. , The spot height h can be determined by the following formula. , Among them B r It is the inner circle of a light spot with radius r.
[0082] For example, a material-related image filter could be a squared norm gradient. This filter could return values as material features related to measurements of the soft and hard transitions and / or roughness of the light spot. Material features can be defined as... .
[0083] For example, a material-dependent image filter can be the standard deviation. The standard deviation of the light spot can be determined by the following formula. , Where µ is the average value given by the following formula: .
[0084] For example, a material-related image filter can be a smoothing filter, such as a Gaussian filter or a median filter. In one embodiment of the smoothing filter, the image filter may refer to an observation that the volume-scattered material exhibits a smaller speckle contrast compared to the diffuse-scattered material. This image filter can quantify the smoothness of the speckle corresponding to the speckle contrast as a material feature. The material feature can be determined by the following formula: , in, This is a smoothing function, such as a median filter or a Gaussian filter. The image filter can include division by a distance z, as described in the formula above. The distance z can be determined, for example, using defocus depth or photon depth ratio techniques and / or by using triangulation techniques. This allows the filter to be insensitive to distance. In one embodiment of the smoothing filter, the smoothing filter can be based on the standard deviation of the extracted speckle noise pattern. The speckle noise pattern N can be described empirically as... , in, It is an image of despeckled light spots. This is the noise term in the model of the speckle pattern. Computing a despeculiar image can be very difficult. Therefore, a despeculiar image can be approximated by a smoothed version of f, i.e. ,in, It is a smoothing operator, such as a Gaussian filter or a median filter. Therefore, an approximation of the speckle pattern can be given by the following equation. .
[0085] The material characteristics of the filter can be determined by the following formula. , Where Var represents the variance function.
[0086] For example, an image filter can be a contrast filter based on the occurrence of gray levels. This material filter can be based on a matrix of gray levels. ,and Let (g1,g2)=[f(x1,y1),f(x2,y2)] be the occurrence rate of gray-scale combinations, and let ρ define the distance between (x1,y1) and (x2,y2), i.e., ρ(x,y)=(x+a,y+b), where a and b are selected from 0 and 1, respectively.
[0087] The material characteristics of a contrast filter based on gray level occurrence can be given by the following formula.
[0088] For example, an image filter could be an energy filter based on the occurrence of gray levels. This material filter is based on the gray level occurrence matrix defined above.
[0089] The material characteristics of an energy filter based on gray level generation can be given by the following formula. .
[0090] For example, an image filter can be a uniform filter based on the occurrence of gray levels. This material filter is based on the gray level occurrence matrix defined above.
[0091] The material characteristics of a uniform filter based on gray level occurrence can be given by the following formula. .
[0092] For example, an image filter can be a dissimilarity filter based on the occurrence of gray levels. This material filter is based on the gray level occurrence matrix defined above.
[0093] The material characteristics of a filter based on the dissimilarity of gray levels can be given by the following formula. .
[0094] For example, an image filter could be a law-based energy filter. This material filter could be based on the law vectors L5=[1,4,6,4,1] and E5=[-1,-2,0,-2,-1] and the matrix L5(E5). T And E5 (L5) T .
[0095] Image f k Convolve with these matrices:
[0096] and .
[0097]
[0098] The material characteristics of the energy filter according to the law can be determined by the following formula: .
[0099] For example, a material-related image filter can be a threshold region filter. This material feature can correlate two regions in the image plane. The first region Ω1 can be a region where the function f is greater than α times the maximum value of f. The second region Ω2 can be a region where the function f is less than α times the maximum value of f but greater than a threshold ε times the maximum value of f. Preferably, α can be 0.5 and ε can be 0.05. Due to speckle or noise, these regions may not simply correspond to the inner and outer circles around the center of the spot. As an example, Ω1 may include speckle or disconnected regions in the outer circle. The material feature can be determined by the following formula. , Where Ω1={x|f(x)>α max(f(x))} and Ω2={x|ε max(f(x)) <f(x)<α max(f(x))}.
[0100] The processing unit can be configured to use material features 2m The material properties of the surface with the generated reflective features are determined by at least one predetermined relationship between the material properties of the surface and the material properties of the surface with the generated reflective features. The predetermined relationship can be one or more of empirical relationships, semi-empirical relationships, and analytically derived relationships. The processing unit may include at least one data storage device for storing the predetermined relationships, such as lookup lists or lookup tables.
[0101] While the feature-based methods described above are sufficient to accurately distinguish skin from surface-scattering materials, differentiating skin from carefully selected attack materials (which also involve volume scattering) is more challenging. Step b) could include the use of artificial intelligence, specifically convolutional neural networks. Using reflection images as input to a convolutional neural network can generate a classification model with sufficient accuracy to distinguish skin from other volume-scattering materials. Since only physically valid information is passed to the network by selecting important regions in the reflection images, a compact training dataset may be required. Furthermore, a very compact network architecture can be generated.
[0102] Specifically, in the skin detection step, at least one parameterized skin classification model can be used. The parameterized skin classification model can be configured to classify skin and other materials using a second image as input. The skin classification model can be parameterized using one or more of machine learning, deep learning, neural networks, or other forms of artificial intelligence. The term "machine learning" as used herein is a broad term and should be given its common and conventional meaning to those skilled in the art, and is not limited to a specific or customized meaning. The term can specifically refer to, but is not limited to, methods of using artificial intelligence (AI) to automatically build models, particularly parameterized models. The term "skin classification model" can refer to a classification model configured to distinguish human skin from other materials. The properties of the skin can be determined by applying an optimization algorithm based on at least one optimization objective on the skin classification model. Machine learning can be based on at least one neural network, particularly a convolutional neural network. The weights and / or topology of the neural network can be predetermined and / or predefined. Specifically, machine learning can be used to train the skin classification model. The skin classification model can include at least one machine learning architecture and model parameters. For example, a machine learning architecture can be or may include one or more of the following: linear regression, logistic regression, random forest, Naive Bayes classification, nearest neighbor, neural network, convolutional neural network, generative adversarial network, support vector machine, or gradient boosting algorithm, etc. As used herein, the term "training" also means learning, is a broad term and will be given its common and conventional meaning to those skilled in the art, and is not limited to a specific or customary meaning. The term may specifically refer to, but is not limited to, the process of building a skin classification model, particularly determining and / or updating the parameters of the skin classification model. The skin classification model may be at least partially data-driven. For example, the skin classification model may be based on experimental data, such as data determined by illuminating multiple people and artificial objects such as masks and recording reflection patterns. For example, training may include using at least one training dataset, wherein the training dataset comprises images of multiple people and artificial objects with known material properties, particularly a second image.
[0103] The skin detection step may include using at least one 2D face and facial landmark detection algorithm configured to provide at least two locations of feature points of a face. For example, the locations could be the eye location, the forehead, or the cheek. The 2D face and facial landmark detection algorithm can provide the locations of feature points of a face, such as the eye location. Because there are subtle differences in reflections in different areas of the face (e.g., the forehead or cheek), region-specific models can be trained. In the skin detection step, at least one region-specific parametric skin classification model is preferably used. The skin classification model may include multiple region-specific parametric skin classification models, for example, for different regions, and / or region-specific data may be used to train the skin classification model, for example, by filtering the images used for training. For example, for training, two different regions may be used, such as the eye-cheek region below the nose, and particularly if sufficient reflective features cannot be identified in that region, the forehead region may be used. However, other regions are also possible.
[0104] If the material property corresponds to at least one property characteristic of skin, the detected face is characterized as skin. The processing unit can be configured to identify reflective features generated by illumination of biological tissue, particularly skin, if its corresponding material property satisfies at least one predetermined or predefined criterion. If the material property indicates "human skin," the reflective feature can be identified as generated by human skin. If the material property is within at least one threshold and / or at least one range, the reflective feature can be identified as generated by human skin. The at least one threshold and / or range can be stored in a table or lookup table and can be determined, for example, empirically, and can be stored as an example in at least one data storage device of the processing unit. The processing unit is configured to otherwise identify the reflective feature as background. Therefore, the processing unit can be configured to assign a material property, such as whether skin is present or not, to each projection point.
[0105] The 3D detection step can be performed after the skin detection step and / or the face detection step. However, other embodiments are also feasible, wherein the 3D detection step is performed before the skin detection step and / or the face detection step. After determining the longitudinal coordinate z in step d), it can be evaluated subsequently. 2m To determine the material properties so that information about the longitudinal coordinate z can be considered in the evaluation. 2m .
[0106] The 3D detection step includes determining second beam profile information for at least four reflection features by analyzing beam profiles of at least four reflection features located within an image region of a second image corresponding to an image region of a first image including the identified geometric features. The 3D detection step may include determining their second beam profile information by analyzing the beam profiles of each of the at least four reflection features. The second beam profile information may include the quotient Q of the region of the beam profile.
[0107] As used herein, the term "beam profile analysis" can generally refer to the evaluation of the beam profile and may include at least one mathematical operation and / or at least one comparison and / or at least symmetry and / or at least one filtering and / or at least one normalization. For example, beam profile analysis may include at least one of histogram analysis steps, calculation of difference measurements, application of neural networks, and application of machine learning algorithms. The processing unit can be configured to symmetry and / or normalize and / or filter the beam profile, particularly removing noise or asymmetry from recordings at larger angles, recording edges, etc. The processing unit can filter the beam profile by removing high spatial frequencies (e.g., through spatial frequency analysis and / or median filtering, etc.). Summarization can be performed by averaging all intensities at the intensity center of the beam spot and at the same distance from the center. The processing unit can be configured to normalize the beam profile to the maximum intensity, particularly taking into account intensity differences due to recording distances. The processing unit can be configured to remove the effects of background light from the beam profile, for example, by imaging in the absence of illumination.
[0108] The processing unit can be configured to determine at least one longitudinal coordinate z of a corresponding reflection feature by analyzing the beam profile of a reflection feature located within an image region of a second image corresponding to an image region of a first image including the identified geometric features. DPR The processing unit can be configured to determine the longitudinal coordinate z of the reflection feature using a technique known as photon depth ratio (also referred to as beam profile analysis). DPR For technical references on depth-to-photon ratio (DPR), see WO 2018 / 091649A1, WO 2018 / 091638A1 and WO 2018 / 091640A1, the entire contents of which are incorporated herein by reference.
[0109] The processing unit can be configured to determine at least one first region and at least one second region of the reflected beam profile of each reflection feature and / or at least one region of interest. The processing unit is configured to integrate the first region and the second region.
[0110] Analysis of the beam profile, one of the reflection characteristics, may include determining at least one first region and at least one second region of the beam profile. The first region of the beam profile may be region A1, and the second region of the beam profile may be region A2. A processing unit may be configured to integrate the first and second regions. The processing unit may be configured to derive a combined signal, particularly the quotient Q, through one or more of the following: dividing the integrated first and second regions, dividing the integrated first and second regions by a multiple, or dividing the integrated first and second regions by a linear combination. The processing unit may be configured to determine at least two regions of the beam profile and / or segment the beam profile into at least two segments comprising different regions of the beam profile, wherein overlap of regions is possible as long as the regions are not identical. For example, the processing unit may be configured to determine multiple regions, such as two, three, four, five, or up to ten regions. The processing unit may be configured to segment the beam spot into at least two regions of the beam profile and / or segment the beam profile into at least two segments comprising different regions of the beam profile. The processing unit may be configured to determine the integral of the beam profile over the corresponding regions for the at least two regions. The processing unit may be configured to compare at least two of the determined integrals. Specifically, the processing unit may be configured to determine at least one first region and at least one second region of the beam profile. As used herein, the term "region of beam profile" generally refers to any region of the beam profile at the location of the optical sensor used to determine the quotient Q. The first region and the second region of the beam profile may be one or both of adjacent or overlapping regions. The first region and the second region of the beam profile may be different in area. For example, the processing unit may be configured to divide the sensor region of a CMOS sensor into at least two sub-regions, wherein the processing unit may be configured to divide the sensor region of the CMOS sensor into at least one left portion and at least one right portion and / or at least one upper portion and at least one lower portion and / or at least one inner portion and at least one outer portion. Alternatively or additionally, the camera may include at least two optical sensors, wherein the photosensitive regions of the first optical sensor and the second optical sensor may be arranged such that the first optical sensor is adapted to determine the first region of the beam profile of the reflection characteristics, and the second optical sensor is adapted to determine the second region of the beam profile of the reflection characteristics. The processing unit may be adapted to integrate the first region and the second region. The processing unit may be configured to determine the longitudinal coordinates using at least one predetermined relationship between the quotient Q and the longitudinal coordinates. The predetermined relation can be one or more of empirical relations, semi-empirical relations, and analytically derived relations. The processing unit may include at least one data storage device for storing predetermined relations, such as lookup lists or lookup tables.
[0111] The first region of the beam profile may substantially include edge information of the beam profile, and the second region of the beam profile may substantially include center information of the beam profile, and / or the first region of the beam profile may substantially include information about the left side of the beam profile, and the second region of the beam profile may substantially include information about the right side of the beam profile. The beam profile may have a center, i.e., the center point of the maximum value of the beam profile and / or the center point of the platform of the beam profile and / or the geometric center of the spot, and a descending edge extending from the center. The second region may include the inner region of the cross-section, and the first region may include the outer region of the cross-section. As used herein, the term "substantially center information" generally refers to a lower proportion of edge information (i.e., a lower proportion of the intensity distribution corresponding to the center) compared to the proportion of center information (i.e., the proportion corresponding to the intensity distribution at the center). Preferably, the center information has a proportion of less than 10% of the edge information, more preferably less than 5%, and most preferably the center information does not include edge content. As used herein, the term "substantially edge information" generally refers to a lower proportion of center information compared to the proportion of edge information. Edge information may include information about the entire beam profile, particularly information from the center and edge regions. Edge information may have a proportion of less than 10% of the center information, preferably less than 5%, and more preferably the edge information does not include center content. If at least one region of the beam profile is close to or surrounds the center and substantially includes the center information, it can be determined and / or selected as a second region of the beam profile. If at least one region of the beam profile includes at least a portion of the descending edge of the cross section, it can be determined and / or selected as a first region of the beam profile. For example, the entire area of the cross section can be determined as the first region.
[0112] Other options for the first region A1 and the second region A2 are also possible. For example, the first region may include a substantially outer region of the beam profile, and the second region may include a substantially inner region of the beam profile. For example, in the case of a two-dimensional beam profile, the beam profile may be divided into a left portion and a right portion, wherein the first region may substantially include the region of the left portion of the beam profile, and the second region may substantially include the region of the right portion of the beam profile.
[0113] Edge information may include information relating to the number of photons in a first region of the beam profile, and center information may include information relating to the number of photons in a second region of the beam profile. The processing unit may be configured to determine the area integral of the beam profile. The processing unit may be configured to determine the edge information by integrating and / or summing over the first region. The processing unit may be configured to determine the center information by integrating and / or summing over the second region. For example, the beam profile may be a trapezoidal beam profile, and the processing unit may be configured to determine the integral of the trapezoid. Furthermore, when a trapezoidal beam profile can be assumed, the determination of the edge and center signals may be replaced by an equivalent evaluation utilizing the properties of the trapezoidal beam profile, such as the determination of the slope and position of the edges and the height of the central plateau, and the edge and center signals may be derived through geometric considerations.
[0114] In one embodiment, A1 may correspond to the entire or complete area of the feature point on the optical sensor. A2 may be the central area of the feature point on the optical sensor. The central area may be a constant value. Compared to the entire area of the feature point, the central area may be smaller. For example, in the case of a circular feature point, the radius of the central area may be 0.1 to 0.9 of the full radius of the feature point, preferably 0.4 to 0.6 of the full radius.
[0115] In one embodiment, the illumination pattern may include at least one line pattern. A1 may correspond to the full linewidth region of the line pattern on the optical sensor, particularly on the photosensitive area of the optical sensor. Compared to the line pattern of the illumination pattern, the line pattern on the optical sensor may be widened and / or shifted, thereby increasing the linewidth on the optical sensor. Specifically, in the case of an optical sensor matrix, the linewidth of the line pattern on the optical sensor may be changed from one column to another. A2 may be the central region of the line pattern on the optical sensor. The linewidth of the central region may be a constant value and may, in particular, correspond to the linewidth in the illumination pattern. The central region may have a smaller linewidth compared to the entire linewidth. For example, the linewidth of the central region may be 0.1 to 0.9 of the full linewidth, preferably 0.4 to 0.6 of the full linewidth. The line pattern may be segmented on the optical sensor. Each column of the optical sensor matrix may include center information of the intensity in the central region of the line pattern and edge information of the intensity from the region extending further outward from the central region of the line pattern to the edge region.
[0116] In one embodiment, the illumination pattern may include at least a dot pattern. A1 may correspond to the entire radius region of the dots in the dot pattern on the optical sensor. A2 may be the central region of the dots in the dot pattern on the optical sensor. The central region may be a constant value. The central region may have a radius compared to the entire radius. For example, the central region may have a radius from 0.1 to 0.9 of the entire radius, preferably from 0.4 to 0.6 of the entire radius.
[0117] The lighting pattern may include both at least one dot pattern and at least one line pattern. Other embodiments besides or alternatives to line and dot patterns are possible.
[0118] The processing unit can be configured to derive the quotient Q by one or more of the following: dividing the first integrated region and the second integrated region, dividing the first integrated region and the second integrated region by a multiple of the first integrated region and the second integrated region, and dividing the first integrated region and the second integrated region by a linear combination.
[0119] The processing unit can be configured to derive the quotient Q by one or more of the following: dividing the first region and the second region, dividing the first region and the second region by a multiple, or dividing the first region and the second region by a linear combination. The processing unit can be configured to derive the quotient Q in the following ways:
[0120] Where x and y are the lateral coordinates, A1 and A2 are the first and second regions of the beam profile, respectively, and E(x, y) represents the beam profile.
[0121] Alternatively, the processing unit may be adapted to determine one or both of center information and edge information from at least one slice or cut of the light spot. This can be achieved, for example, by replacing the region integral in the quotient Q with a line integral along the slice or cut. To improve accuracy, multiple slices or cuts passing through the light spot can be used and averaged. In the case of an elliptical light spot profile, averaging multiple slices or cuts may result in improved distance information.
[0122] For example, in the case of an optical sensor with a pixel matrix, the processing unit can be configured to evaluate the beam profile in the following way: - Identify the pixel with the highest sensor signal and that forms at least one center signal; - Evaluate the sensor signals of the matrix and form at least one sum signal; - Determine the quotient Q by combining the center signal and the sum signal; and - Determine at least one longitudinal coordinate z of the object by evaluating the quotient Q.
[0123] As used herein, "sensor signal" generally refers to a signal generated by an optical sensor and / or at least one pixel of the optical sensor in response to illumination. Specifically, a sensor signal can be or may include at least one electrical signal, such as at least one analog electrical signal and / or at least one digital electrical signal. More specifically, a sensor signal can be or may include at least one voltage signal and / or at least one current signal. More specifically, a sensor signal may include at least one photocurrent. Furthermore, the original sensor signal can be used, or an auxiliary sensor signal can be generated by processing or preprocessing the sensor signal using a display device, optical sensor, or any other element, which can also be used as a sensor signal, for example, by preprocessing through filtering, etc. The term "center signal" generally refers to at least one sensor signal that substantially includes center information of the beam profile. As used herein, the term "maximum sensor signal" refers to a local maximum value or one or both of the maximum values in the region of interest. For example, the center signal can be the signal of the pixel with the highest sensor signal among multiple sensor signals generated by pixels of the entire matrix or the region of interest within the matrix, wherein the region of interest can be predetermined or determinable within the image generated by the pixels of the matrix. The center signal can originate from a single pixel or a group of optical sensors. In the latter case, as an example, the sensor signals of a group of pixels can be added, integrated, or averaged to determine the center signal. The group of pixels generating the center signal can be a group of adjacent pixels, such as pixels located less than a predetermined distance from the actual pixel with the highest sensor signal, or it can be a group of pixels generating sensor signals within a predetermined range from the highest sensor signal. The group of pixels generating the center signal can be selected as large as possible to allow for the maximum dynamic range. The processing unit can be adapted to determine the center signal by integrating multiple sensor signals, such as multiple pixels surrounding the pixel with the highest sensor signal. For example, the beam profile can be a trapezoidal beam profile, and the processing unit can be adapted to determine the integral of the trapezoid, particularly the integral of the trapezoidal plateau.
[0124] As described above, the center signal can typically be a single sensor signal, such as a sensor signal from the pixel at the center of the spot, or it can be a combination of multiple sensor signals, such as a combination of sensor signals from the pixel at the center of the spot, or an auxiliary sensor signal derived by processing sensor signals derived from one or more of the above possibilities. The determination of the center signal can be performed electronically, as the comparison of sensor signals is quite simple with conventional electronic equipment, or it can be performed entirely or partially by software. Specifically, the center signal can be selected from the following group: the highest sensor signal; the average value of a group of sensor signals within a predetermined tolerance range from the highest sensor signal; the average value of sensor signals from a group of pixels including the pixel with the highest sensor signal and a predetermined group of adjacent pixels; the sum of sensor signals from a group of pixels including the pixel with the highest sensor signal and a predetermined group of adjacent pixels; the sum of a group of sensor signals within a predetermined tolerance range from the highest sensor signal; the average value of a group of sensor signals above a predetermined threshold; the sum of a group of sensor signals above a predetermined threshold; the integral of sensor signals from a group of optical sensors including the optical sensor with the highest sensor signal and a predetermined group of adjacent pixels; the integral of a group of sensor signals within a predetermined tolerance range from the highest sensor signal; and the integral of a group of sensor signals above a predetermined threshold.
[0125] Similarly, the term "summary signal" generally refers to a signal that essentially includes edge information of the beam profile. For example, a summary signal can be derived by adding sensor signals, integrating sensor signals, or averaging sensor signals over the entire matrix or a region of interest within the matrix, where the region of interest can be predetermined or defined within an image generated by the optical sensors of the matrix. When adding, integrating, or averaging sensor signals, the actual optical sensors that generated the sensor signals can be excluded from the addition, integration, or averaging, or alternatively, can be included in the addition, integration, or averaging. The processing unit can be adapted to determine the summary signal by integrating the signal over the entire matrix or the region of interest within the matrix. For example, the beam profile can be a trapezoidal beam profile, and the processing unit can be adapted to determine the integral over the entire trapezoid. Furthermore, when a trapezoidal beam profile can be assumed, the determination of edge and center signals can be replaced by an equivalent evaluation utilizing the properties of the trapezoidal beam profile, such as the determination of the slope and position of the edges and the height of the central plateau, and the derivation of edge and center signals through geometric considerations.
[0126] Similarly, center and edge signals can be determined using segments of the beam profile (e.g., circular segments of the beam profile). For example, the beam profile can be divided into two segments by a secant or chord that does not pass through the center of the beam profile. Thus, one segment will essentially contain edge information, while the other segment will essentially contain center information. For example, to further reduce the amount of edge information in the center signal, the edge signal can be subtracted from the center signal.
[0127] The quotient Q can be a signal generated by combining a center signal and a sum signal. Specifically, the determination can include one or more of the following: forming a quotient of the center signal and the sum signal, or vice versa; forming a quotient of a multiple of the center signal and a multiple of the sum signal, or vice versa; forming a quotient of a linear combination of the center signals and a linear combination of the sum signals, or vice versa. Alternatively or additionally, the quotient Q can include any signal or combination of signals containing at least one piece of information regarding the comparison between the center signal and the sum signal.
[0128] As used herein, the term "vertical coordinate of a reflection feature" refers to the distance between an optical sensor and a point in the scene emitting the corresponding illumination feature. The processing unit may be configured to determine the vertical coordinate using at least one predetermined relationship between the quotient Q and the vertical coordinate. The predetermined relationship may be one or more of empirical, semi-empirical, and analytically derived relationships. The processing unit may include at least one data storage device for storing the predetermined relationship, such as a lookup list or lookup table.
[0129] The processing unit can be configured to execute at least one photon depth ratio algorithm that calculates the distance with all reflection features of zeroth order and higher.
[0130] The 3D detection step may include determining at least one depth level from second beam profile information of the reflected feature using a processing unit.
[0131] The processing unit can be configured to determine a depth map of at least a portion of a scene by determining at least one depth information of reflective features located within an image region of a second image corresponding to an image region of a first image including the identified geometric features. As used herein, the term "depth" or depth information may refer to the distance between an object and an optical sensor and may be given by a longitudinal coordinate. As used herein, the term "depth map" may refer to the spatial distribution of depth. The processing unit can be configured to determine the depth information of reflective features using one or more of the following techniques: photon depth ratio, structured light, beam profile analysis, time-of-flight, shape-from-motion, depth-from-focus, triangulation, depth-from-defocus, and stereo sensor. The depth map may be a sparsely filled depth map comprising a few entries. Alternatively, the depth may be crowded, comprising a large number of entries.
[0132] If the depth level deviates from a predetermined or predefined depth level of the planar object, the detected face is characterized as a 3D object. Step c) may include using 3D topological data of the face in front of the camera. The method may include determining curvature based on at least four reflection features located within an image region of a second image corresponding to an image region of a first image including the identified geometric features. The method may include comparing the curvature determined from the at least four reflection features with a predetermined or predefined depth level of the planar object. If the curvature exceeds the assumed curvature of the planar object, the detected face may be characterized as a 3D object; otherwise, it is characterized as a planar object. The predetermined or predefined depth level of the planar object may be stored in at least one data memory of the processing unit, such as a lookup list or lookup table. The predetermined or predefined depth level of the planar object may be determined experimentally and / or may be a theoretical level of the planar object. The predetermined or predefined depth level of the planar object may be at least one limit of at least one curvature and / or a range of at least one curvature.
[0133] The 3D features identified in step c) allow for the differentiation between high-quality photographs and 3D facial structures. The combination of steps b) and c) can enhance the reliability of authentication against attacks. 3D features can be combined with material features to improve security levels. Since the same computational pipeline can be used to generate the input data for skin classification and the 3D point cloud, both properties can be computed from the same frame with relatively low computational cost.
[0134] Preferably, an authentication step can be performed after steps a) to c). The authentication step can be performed partially after each of steps a) to c). Authentication can be aborted if no face is detected in step a) and / or if it is determined in step b) that the reflection features are not generated by skin and / or the depth map involves a planar object in step c). The authentication step includes authenticating the detected face by using at least one authentication unit if the face detected in step b) is characterized as skin and the face detected in step c) is characterized as a 3D object.
[0135] Steps a) through d) can be performed using at least one device, such as at least one mobile device, like a mobile phone, smartphone, etc., where access to the device is protected by using facial authentication. Other devices are also possible, such as access control devices that control access to buildings, machines, cars, etc. The method may include allowing access to the device if the detected face is authenticated.
[0136] The method may include at least one enrollment step. In the enrollment step, a user of the device may be registered. As used herein, the term "enrollment" may refer to the process of registering and / or signing up and / or instructing a user to subsequently use the device. Typically, enrollment may be performed upon first use of the device and / or startup of the device. However, embodiments in which multiple users may be registered consecutively, making it feasible to perform and / or repeat enrollment at any time during device use, are possible. Enrollment may include generating a user account and / or user profile. Enrollment may include inputting and storing user data, particularly image data, via at least one user interface. Specifically, at least one 2D image of the user is stored in at least one database. The enrollment step may include imaging at least one image of the user, particularly multiple images. Images may be recorded from different orientations and / or the user may change their orientation. Additionally, the enrollment step may include generating at least one 3D image and / or depth map for the user, which may be used for comparison in step d). The database may be a database of the device, such as the database of a processing unit, and / or may be an external database such as the cloud. The method includes identifying the user by comparing the user's 2D image with a first image. The method according to the invention can allow for a significant improvement in the presentation attack detection capability of biometric authentication methods. To improve overall authentication, in addition to the user's 2D image, personal material fingerprints and 3D topological features can be stored during the registration process. This allows for multi-factor authentication within a single device using 2D, 3D, and material-derived features.
[0137] The method using beam profile analysis technology according to the present invention provides a concept for reliably detecting human skin by analyzing the reflection of laser spots on the face (particularly in the NIR range) and distinguishing it from reflections from attack materials produced by mimicking the face. Furthermore, beam profile analysis simultaneously provides depth information by analyzing the same camera frame. Therefore, 3D and skin-safe features can be provided using the exact same technology.
[0138] Since 2D images of the face can also be recorded simply by turning off the laser illumination, a completely secure facial recognition pipeline can be established to solve the above problems.
[0139] As the laser wavelength shifts into the NIR region, the reflectivity of human skin of different skin tones becomes more similar. The differences are minimal at a wavelength of 940 nm. Therefore, skin color does not contribute to skin authentication.
[0140] It may not be necessary to perform time-consuming analysis on a series of frames, since the presentation attack detection (via skin classification) is provided by only one frame. The time frame used to perform the full method can be ≤500 ms, preferably ≤250 ms. However, embodiments in which multiple frames can be used to perform skin detection may be feasible. Depending on the confidence level of identifying reflection features in the second image and the speed of the method, the method may include sampling reflection features on multiple frames to achieve a more stable classification.
[0141] Besides accuracy, execution speed and power consumption are also important requirements. For security reasons, further restrictions on the availability of computing resources can be introduced. For example, steps a) to d) can be run in a secure region of the processing unit to avoid any software-based manipulation during program execution. The compact nature of the aforementioned material detection network can address this issue by exhibiting excellent runtime behavior in the secure region, whereas traditional PAD solutions require checking several consecutive frames, resulting in high computational cost and long response time.
[0142] In another aspect of the invention, a computer program for face authentication is configured to cause the computer or computer network to perform the method according to the invention, fully or partially, when executed on a computer or computer network, wherein the computer program is configured to perform and / or perform at least steps a) to d) of the method according to the invention. Specifically, the computer program may be stored on a computer-readable data carrier and / or a computer-readable storage medium.
[0143] As used herein, the terms "computer-readable data carrier" and "computer-readable storage medium" can specifically refer to a non-transitory data storage device, such as a hardware storage medium on which computer-executable instructions are stored. A computer-readable data carrier or storage medium can specifically be or can include storage media such as random access memory (RAM) and / or read-only memory (ROM).
[0144] Therefore, specifically, one, more, or even all of the above method steps can be executed using a computer or computer network, preferably using a computer program.
[0145] On the other hand, the computer-readable storage medium includes instructions that, when executed by a computer or computer network, cause at least steps a) to d) of the method according to the invention to be performed.
[0146] This document further discloses and proposes a data carrier having a data structure stored thereon, which, after being loaded into a computer or computer network, such as into the working memory or main memory of the computer or computer network, can perform methods according to one or more embodiments disclosed herein.
[0147] This document also discloses and proposes a computer program product having program code means stored on a machine-readable medium for performing methods according to one or more embodiments disclosed herein when the program is executed on a computer or computer network. As used herein, a computer program product refers to a program that is a tradable product. The product can generally exist in any format, such as in paper format, or on a computer-readable data carrier and / or computer-readable storage medium. Specifically, the computer program product can be distributed via a data network.
[0148] Finally, this document discloses and proposes a modulated data signal containing computer system or computer network readable instructions for performing methods according to one or more embodiments disclosed herein.
[0149] Referring to the computer implementation aspects of the present invention, one or more, or even all, method steps of the methods according to one or more embodiments disclosed herein can be performed using a computer or computer network. Therefore, generally, any method step, including the provision and / or manipulation of data, can be performed using a computer or computer network. In general, these method steps can include any method steps, typically except for those requiring manual operation.
[0150] Specifically, the present invention further discloses: - A computer or computer network, including at least one processor, wherein the processor is adapted to perform a method according to one of the embodiments described in this specification. - A computer-loadable data structure suitable for performing a method according to one of the embodiments described in this specification when the data structure is executed on a computer. - A computer program, wherein the computer program is adapted to perform, when executed on a computer, a method according to one of the embodiments described in this specification. - A computer program, including program means, for performing a method according to one of the embodiments described in this specification when the computer program is executed on a computer or a computer network. - A computer program, including a program means according to the foregoing embodiments, wherein the program means is stored on a computer-readable storage medium. - A storage medium, wherein a data structure is stored on the storage medium, and wherein the data structure is adapted to execute a method according to one of the embodiments described herein after being loaded into the main memory and / or working memory of a computer or computer network. - A computer program product having program code means, wherein the program code means may be stored or stored on a storage medium, and if the program code means is executed on a computer or computer network, it is used to perform a method according to one of the embodiments described in this specification.
[0151] In another aspect, a mobile device is disclosed, comprising at least one camera, at least one illumination unit, and at least one processing unit. The mobile device is configured to perform at least steps a) to c) of the face authentication method according to the invention, and optionally step d). Step d) can be performed using at least one authentication unit. The authentication unit can be a unit of the mobile device or an external authentication unit. For definitions and embodiments of the mobile device, refer to the definitions and embodiments described with respect to the method.
[0152] In another aspect of the invention, for the purpose of detecting biological presentation attacks, a method according to the invention is proposed, such as one or more embodiments given above or further detailed below.
[0153] In general, the following embodiments are considered preferred in the context of this invention: Example 1. A method for facial authentication, comprising the following steps: a) at least one face detection step, wherein the face detection step includes: using at least one camera to determine at least one first image, wherein the first image includes at least one two-dimensional image of a scene suspected of including a face, wherein the face detection step includes: using at least one processing unit to detect the face in the first image by recognizing at least one predefined or predetermined geometric feature characteristic for the face in the first image; b) At least one skin detection step, wherein the skin detection step includes: using at least one illumination unit to project at least one illumination pattern including a plurality of illumination features onto the scene, and using at least one camera to determine at least one second image, wherein the second image includes a plurality of reflection features generated by the scene in response to illumination of the illumination features, wherein each of the reflection features includes at least one beam profile, wherein the skin detection step includes: using the processing unit to determine first beam profile information of the at least one reflection feature by analyzing the beam profile of at least one reflection feature located in an image region of the second image corresponding to an image region of the first image including the identified geometric features, and determining at least one material property of the reflection feature from the first beam profile information, wherein if the material property corresponds to at least one property characteristic for skin, the detected face is characterized as skin; c) At least one 3D detection step, wherein the 3D detection step includes: using the processing unit to determine second beam profile information of at least four reflection features by analyzing beam profiles of at least four reflection features located in the image region of the second image corresponding to the image region of the first image including the identified geometric features; and determining at least one depth level from the second beam profile information of the reflection features, wherein if the depth level deviates from a predetermined or predefined depth level of a planar object, the detected face is characterized as a 3D object; d) At least one authentication step, wherein the authentication step includes: authenticating the detected face by using at least one authentication unit if the face detected in step b) is characterized as skin and the face detected in step c) is characterized as a 3D object.
[0154] Example 2. The method according to the foregoing embodiments, wherein steps a) to d) are performed by using at least one device, wherein access to the device is protected by using facial authentication, wherein the method includes: allowing access to the device if the detected face is authenticated.
[0155] Example 3. The method according to the foregoing embodiments, wherein the method includes at least one registration step, wherein in the registration step, a user of the device is registered, wherein at least one 2D image of the user is stored in at least one database, wherein the method includes: identifying the user by comparing the 2D image of the user with the first image.
[0156] Example 4. The method according to any of the foregoing embodiments, wherein, in the skin detection step, at least one parameterized skin classification model is used, wherein the parameterized skin classification model is configured to classify skin and other materials by using the second image as input.
[0157] Example 5. The method according to the foregoing embodiments, wherein the skin classification model is parameterized by using machine learning, wherein the property characteristics of the skin are determined by applying an optimization algorithm on the skin classification model in terms of at least one optimization objective.
[0158] Example 6. The method according to any one of the foregoing two embodiments, wherein the skin detection step includes: using at least one 2D face and facial landmark detection algorithm, the 2D face and facial landmark detection algorithm being configured to provide at least two locations of feature points of a face, wherein, in the skin detection step, at least one region-specific parametric skin classification model is used.
[0159] Example 7. The method according to any of the foregoing embodiments, wherein the illumination pattern comprises a periodic grid of laser spots.
[0160] Example 8. The method according to any of the foregoing embodiments, wherein the illumination feature has a wavelength in the near-infrared (NIR) range.
[0161] Example 9. The method according to the foregoing embodiments, wherein the illumination feature has a wavelength of 940 nm.
[0162] Example 10. The method according to any of the foregoing embodiments, wherein a plurality of second images are determined, wherein the reflection features of the plurality of second images are used for skin detection in step b) and / or for 3D detection in step c).
[0163] Example 11. The method according to any of the foregoing embodiments, wherein the camera is or includes at least one near-infrared camera.
[0164] Example 12. A computer program for facial authentication, configured to, when executed on a computer or computer network, cause the computer or computer network to perform, fully or partially, the method according to any of the foregoing embodiments, wherein the computer program is configured to perform and / or implement at least steps a) to d) of the method according to any of the foregoing embodiments.
[0165] Example 13. A computer-readable storage medium including instructions that, when executed by a computer or computer network, cause at least steps a) to d) of the method according to any of the foregoing embodiments relating to the method to be performed.
[0166] Example 14. A mobile device including at least one camera, at least one illumination unit, and at least one processing unit, the mobile device being configured to perform at least steps a) to c) of a method for facial authentication according to any of the foregoing embodiments of the method, and optionally step d).
[0167] Example 15. The use of the method according to any one of the foregoing embodiments for biological presentation attack detection. Attached Figure Description
[0168] Further optional details and features of the invention will be apparent from the description of preferred exemplary embodiments in conjunction with the dependent claims. Specific features may be implemented individually or in combination with other features herein. The invention is not limited to exemplary embodiments. Exemplary embodiments are schematically illustrated in the accompanying drawings. The same reference numerals in the various drawings denote the same elements or elements having the same function, or elements corresponding to each other in terms of their function.
[0169] Specifically, in the diagram: Figure 1 An embodiment of a method for facial authentication according to the present invention is shown; Figure 2 An embodiment of a mobile device according to the present invention is shown; and Figure 3 The experimental results are shown. Detailed Implementation
[0170] Figure 1A flowchart of a method for facial authentication according to the present invention is shown. Facial authentication may include verifying an identified object or a portion thereof as a face. Specifically, authentication may include distinguishing a real face from attack material generated to mimic a face. Authentication may include verifying the identity of the corresponding user and / or assigning an identity to the user. Authentication may include generating and / or providing identity information, for example to other devices, such as to at least one authorized device, for authorizing access to mobile devices, machines, vehicles, buildings, etc. The identity information can be proven through authentication. For example, the identity information may be and / or may include at least one identity token. In the case of successful authentication, the identified object or a portion thereof is verified as a real face and / or the identity of the object, particularly the user, is verified.
[0171] The method includes the following steps: a) (Ref. 110) At least one face detection step, wherein the face detection step includes determining at least one first image by using at least one camera 112, wherein the first image includes at least one two-dimensional image of a scene suspected of including a face, wherein the face detection step includes detecting a face in the first image by using at least one processing unit 114 to identify at least one predefined or predetermined geometric feature of a face in the first image. (b) (reference numeral 116) at least one skin detection step, wherein the skin detection step includes projecting at least one illumination pattern comprising a plurality of illumination features onto a scene using at least one illumination unit 118 and determining at least one second image using at least one camera 112, wherein the second image includes a plurality of reflection features generated by the scene in response to illumination of the illumination features, wherein each reflection feature includes at least one beam profile, wherein the skin detection step includes determining first beam profile information of at least one reflection feature by analyzing the beam profile of at least one reflection feature among the reflection features located in an image region of the second image, the image region corresponding to an image region of the first image including identified geometric features, and determining at least one material property of the reflection feature from the first beam profile information using a processing unit 114, wherein if the material property corresponds to at least one property characteristic of skin, the detected face is characterized as skin; c) (Ref. 120) at least one 3D detection step, wherein the 3D detection step determines second beam profile information of at least four reflective features by analyzing beam profiles of at least four reflective features in an image region located in a second image, the image region corresponding to an image region of the first image including the identified geometric features, wherein if the depth level deviates from a predetermined or predefined depth level of a planar object, the detected face is characterized as a 3D object. d) (reference numeral 122) at least one authentication step, wherein the authentication step includes authenticating the detected face by using at least one authentication unit if the face detected in step b) is characterized as skin and the face detected in step c) is characterized as a 3D object.
[0172] The method steps can be executed in a given order or in a different order. Furthermore, there may be one or more additional method steps not listed. Moreover, one, more, or even all of the method steps can be executed repeatedly.
[0173] Camera 112 may include at least one imaging element configured to record or capture spatially resolved one-dimensional, two-dimensional, or even three-dimensional optical data or information. Camera 112 may be a digital camera. As an example, camera 112 may include at least one camera chip, such as at least one CCD chip and / or at least one CMOS chip configured to record images. Camera 112 may be or may include at least one near-infrared camera. Images may involve data recorded using camera 112, such as multiple electronic readings from an imaging device, such as pixels of the camera chip. In addition to at least one camera chip or imaging chip, camera 112 may also include additional elements, such as one or more optical elements, such as one or more lenses. As an example, camera 112 may be a fixed-focus camera with at least one lens that is fixedly adjusted relative to the camera. However, alternatively, camera 112 may also include one or more variable lenses that can be adjusted automatically or manually.
[0174] Camera 112 can be a camera of mobile device 124, such as a camera of a laptop, tablet, or, in particular, a cellular phone such as a smartphone. Therefore, specifically, camera 112 can be part of mobile device 124, which, in addition to at least one camera 112, also includes one or more data processing devices, such as one or more data processors. However, other cameras are also possible. Mobile device 124 can be a mobile electronic device, more specifically a mobile communication device such as a cellular phone or smartphone. Alternatively or additionally, mobile device 124 can also refer to a tablet computer or another type of portable computer. Figure 2 An embodiment of a mobile device according to the present invention is shown.
[0175] Specifically, camera 112 may be or may include at least one optical sensor 126 having at least one photosensitive region. Optical sensor 126 may specifically be or may include at least one photodetector, preferably an inorganic photodetector, more preferably an inorganic semiconductor photodetector, and most preferably a silicon photodetector. Specifically, optical sensor 126 may be sensitive in the infrared spectral range. Optical sensor 126 may include at least one sensor element comprising a pixel matrix. All pixels of the matrix or at least one group of optical sensors in the matrix may specifically be identical. Specifically, groups of identical pixels in the matrix may be provided for different spectral ranges, or all pixels may be identical in terms of spectral sensitivity. Furthermore, the pixel size and / or their electronic or photoelectric properties may be identical. Specifically, optical sensor 126 may be or may include at least one array of inorganic photodiodes sensitive in the infrared spectral range, preferably sensitive in the range of 700 nm to 3.0 micrometers. Specifically, optical sensor 126 may be sensitive in a portion of the near-infrared region suitable for silicon photodiodes, specifically in the range of 700 nm to 1100 nm. The infrared optical sensor that can be used in the optical sensor can be a commercially available infrared optical sensor, such as the trinamiX from D-67056, Ludwigshafen ait Rhine, Germany. TM GmbH with Hertzstueck TM Commercially available infrared optical sensors. Therefore, as an example, optical sensor 126 may include at least one intrinsic photovoltaic optical sensor, more preferably at least one semiconductor photodiode selected from the group consisting of: Ge photodiodes, InGaAs photodiodes, extended InGaAs photodiodes, InAs photodiodes, InSb photodiodes, and HgCdTe photodiodes. Alternatively, the optical sensor may include at least one extrinsic photovoltaic optical sensor, more preferably at least one semiconductor photodiode selected from the group consisting of: Ge:Au photodiodes, Ge:Hg photodiodes, Ge:Cu photodiodes, Ge:Zn photodiodes, Si:Ga photodiodes, and Si:As photodiodes. Alternatively, optical sensor 126 may include at least one photoconductivity sensor, such as a PbS or PbSe sensor, or a thermal calorimeter, preferably selected from the group consisting of V0 thermal calorimeters and amorphous Si thermal calorimeters.
[0176] Specifically, the optical sensor 126 can be sensitive in the near-infrared region. Specifically, the optical sensor 126 can be sensitive in a portion of the near-infrared region suitable for silicon photodiodes, specifically in the range of 700 nm to 1000 nm. Specifically, the optical sensor 126 can be sensitive in the infrared spectral range, specifically in the range of 780 nm to 3.0 micrometers. For example, the optical sensor 126 can be or can include at least one element selected from the group consisting of: CCD sensor elements, CMOS sensor elements, photodiodes, photovoltaic cells, photoconductors, phototransistors, or any combination thereof. Any other type of photosensitive element can be used. Photosensitive elements can generally be made wholly or partially of inorganic materials and / or can be made wholly or partially of organic materials. Most commonly, one or more photodiodes can be used, such as commercially available photodiodes, such as inorganic semiconductor photodiodes.
[0177] Camera 112 may also include at least one transmission device (not shown here). Camera 112 may include at least one optical element selected from the group consisting of: a transmission device, such as at least one lens and / or at least one lens system, at least one diffractive optical element. The transmission device may be adapted to guide the light beam onto optical sensor 126. Specifically, the transmission device may include one or more of the following: at least one lens, such as at least one lens selected from the group consisting of: at least one focusable lens, at least one aspherical lens, at least one spherical lens, at least one Fresnel lens; at least one diffractive optical element; at least one concave mirror; at least one beam deflection element, preferably at least one mirror; at least one beam splitter element, preferably at least one of a beam splitter cube or beam splitter; at least one multi-lens system. The transmission device may have a focal length. Thus, the focal length constitutes a measure of the transmission device's ability to converge an incident light beam. Therefore, the transmission device may include one or more imaging elements that can have the effect of a converging lens. For example, the transmission device may have one or more lenses, particularly one or more refractive lenses, and / or one or more convex mirrors. In this example, the focal length may be defined as the distance from the center of the thin refractive lens to the principal focal point of the thin lens. For converging thin refractive lenses, such as convex or biconvex thin lenses, the focal length can be considered positive and can provide a distance at which the collimated beam illuminating the thin lens, acting as a transmission device, can be focused into a single spot. Additionally, the transmission device may include at least one wavelength selection element, such as at least one filter. Furthermore, the transmission device can be designed to apply a predefined beam profile to electromagnetic radiation (e.g., at the sensor region and particularly at the sensor region's location). In principle, the above-described alternative embodiments of the transmission device can be implemented individually or in any desired combination.
[0178] The transmission device may have an optical axis. The transmission device may constitute a coordinate system, where the longitudinal coordinate is the coordinate along the optical axis, and where d is the spatial offset relative to the optical axis. The coordinate system may be a polar coordinate system, where the optical axis of the transmission device forms a z-axis, and where the distance from the z-axis and the polar angle can be used as additional coordinates. Directions parallel to or antiparallel to the z-axis may be considered longitudinal directions, and coordinates along the z-axis may be considered longitudinal coordinates. Any direction perpendicular to the z-axis may be considered transverse directions, and polar coordinates and / or polar angles may be considered transverse coordinates.
[0179] Camera 112 is configured to determine at least one image of a scene, particularly a first image. The scene may refer to a spatial region. The scene may include a face and its surrounding environment during authentication. The first image itself may include pixels, which are related to the pixels of a matrix of sensor elements. The first image is at least one two-dimensional image having information about lateral coordinates, such as dimensions of height and width.
[0180] The face detection step 110 includes detecting a face in the first image by recognizing at least one predefined or predetermined geometric feature of the face in the first image using at least one processing unit 114. As an example, at least one processing unit 114 may include software code stored thereon, the software code comprising a plurality of computer commands. The processing unit 114 may provide one or more hardware elements for performing one or more specified operations and / or may provide one or more processors on which software for performing one or more specified operations runs. Operations including evaluating the image may be performed by at least one processing unit 114. Thus, as an example, one or more instructions may be implemented in software and / or hardware. Thus, as an example, the processing unit 114 may include one or more programmable devices, such as one or more computers, application-specific integrated circuits (ASICs), digital signal processors (DSPs), or field-programmable gate arrays (FPGAs), configured to perform the aforementioned evaluation. However, additionally or alternatively, the processing unit may also be implemented entirely or partially in hardware. The processing unit 114 and the camera 112 may be fully or partially integrated into a single device. Thus, typically, the processing unit 114 may also be part of the camera 112. Alternatively, the processing unit 114 and the camera 112 may be fully or partially embodied as separate devices.
[0181] Detecting a face in the first image may include identifying at least one predefined or predetermined geometric feature of the face. The geometric feature of the face may be at least one geometry-based feature that describes the shape of the face and its components, particularly one or more of the nose, eyes, mouth, or eyebrows. Processing unit 114 may include at least one database in which the geometric feature of the face is stored, such as in a lookup table. Techniques for identifying at least one predefined or predetermined geometric feature of a face are generally known to those skilled in the art. For example, face detection may be performed as described in Masi, Lacopo et al., “Deep face recognition: A survey,” 31st SIBGRAPI Conference on Graphics, Patterns and Images (SIBGRAPI), IEEE, 2018, the entire contents of which are incorporated herein by reference.
[0182] Processing unit 114 can be configured to perform at least one image analysis and / or image processing to identify geometric features. Image analysis and / or image processing may use at least one feature detection algorithm. Image analysis and / or image processing may include one or more of the following: filtering; selecting at least one region of interest; background correction; decomposition into color channels; decomposition into hue, saturation, and / or brightness channels; frequency decomposition; singular value decomposition; applying a speckle detector; applying an angle detector; applying a determinant of a Hessian filter; applying a region detector based on principle curvature; applying a gradient position and orientation histogram algorithm; applying a histogram of orientation gradient descriptors; applying an edge detector; applying a differential edge detector; applying a Canny edge detector; applying a Laplacian Gaussian filter; applying a difference Gaussian filter; applying a Sobel operator; applying a Laplacian operator; applying a Scharr operator; applying a Prewitt operator; applying a Roberts operator; applying a Kirsch operator; applying a high-pass filter; applying a low-pass filter; applying a Fourier transform; applying a Radon transform; applying a Hough transform; applying a wavelet transform; thresholding; creating a binary image. The region of interest can be determined manually by the user or automatically, for example, by identifying features within the first image.
[0183] Specifically, after the face detection step 110, a skin detection step 116 can be performed, including projecting at least one lighting pattern comprising multiple lighting features onto the scene using at least one lighting unit 118. However, embodiments in which the skin detection step 116 is performed before the face detection step 110 are feasible.
[0184] The lighting unit 118 can be configured to provide a lighting pattern for scene illumination. The lighting unit 118 can be adapted to directly or indirectly illuminate the scene, wherein the lighting pattern is affected by the surface of the scene, particularly by reflection or scattering, and thereby at least partially directed towards the camera. The lighting unit 118 can be configured to illuminate the scene, for example, by directing a light beam to the scene and reflecting that beam. The lighting unit 118 can be configured to generate an illumination beam for illuminating the scene.
[0185] The illumination unit 118 may include at least one light source. The illumination unit 118 may include multiple light sources. The illumination unit 118 may include artificial lighting sources, particularly at least one laser source and / or at least one incandescent lamp and / or at least one semiconductor light source, such as at least one light-emitting diode, particularly organic and / or inorganic light-emitting diodes. The illumination unit 118 may be configured to generate at least one illumination pattern in the infrared region. The illumination feature may have a wavelength in the near-infrared (NIR) range. The illumination feature may have a wavelength of approximately 940 nm. At this wavelength, melanin absorption is depleted, so the dark and bright complexes reflect almost the same light. However, other wavelengths in the NIR region are also possible, such as one or more of 805 nm, 830 nm, 835 nm, 850 nm, 905 nm, or 980 nm. Furthermore, using light in the near-infrared region makes the light undetectable to the human eye or only weakly detectable by the human eye, but still detectable by silicon sensors, particularly standard silicon sensors.
[0186] The illumination unit 118 may be or may include at least one multi-beam source. For example, the illumination unit 118 may include at least one laser source and one or more diffractive optical elements (DOEs). Specifically, the illumination unit 118 may include at least one laser and / or laser source. Various types of lasers may be employed, such as semiconductor lasers, dual heterostructure lasers, external cavity lasers, split-confined heterostructure lasers, quantum cascade lasers, distributed Bragg reflector lasers, polaron lasers, hybrid silicon lasers, extended cavity diode lasers, quantum dot lasers, bulk Bragg grating lasers, indium arsenide lasers, transistor lasers, diode-pumped lasers, distributed feedback lasers, quantum well lasers, interband cascade lasers, gallium arsenide lasers, semiconductor ring lasers, extended cavity diode lasers, or vertical cavity surface-emitting lasers. Alternatively or additionally, non-laser source materials, such as LEDs and / or bulbs, may be used. The illumination unit 118 may include one or more diffractive optical elements (DOEs) suitable for generating illumination patterns. For example, the illumination unit 118 may be adapted to generate and / or project point clouds. The illumination unit 118 may include one or more of the following: at least one digital light processing projector, at least one LCoS projector, at least one spatial light modulator; at least one diffractive optical element; at least one light-emitting diode array; at least one laser source array. Using at least one laser source as the illumination unit 118 is particularly preferred due to its typically defined beam profile and other properties of operability. The illumination unit 118 may be integrated into the housing of the camera 112 or may be separate from the camera 112.
[0187] A lighting pattern includes at least one lighting feature suitable for at least a portion of a lighting scene. A lighting pattern may include a single lighting feature. A lighting pattern may include multiple lighting features. A lighting pattern may be selected from the group consisting of: at least one dot pattern; at least one line pattern; at least one stripe pattern; at least one checkerboard pattern; at least one pattern including an arrangement of periodic or aperiodic features. A lighting pattern may include regular and / or constant and / or periodic patterns, such as triangular patterns, rectangular patterns, hexagonal patterns, or patterns including additional convex joins. A lighting pattern may present at least one lighting feature selected from the group consisting of: at least one dot; at least one row; at least two lines, such as parallel or intersecting lines; at least one dot and one line; at least one arrangement of periodic or aperiodic features; at least one feature of arbitrary shape. The illumination pattern may include at least one pattern selected from the group consisting of: at least one point pattern, particularly a pseudo-random point pattern; a random point pattern or a quasi-random pattern; at least one Sobol pattern; at least one quasi-periodic pattern; at least one pattern including at least one known feature; at least one regular pattern; at least one triangular pattern; at least one hexagonal pattern; at least one rectangular pattern, including at least one pattern with convex uniform tiling; at least one line pattern including at least one line; at least one line pattern including at least two lines (e.g., parallel or intersecting lines). For example, illumination unit 118 may be adapted to generate and / or project a point cloud. Illumination unit 118 may include at least one light projector adapted to generate a point cloud such that the illumination pattern may include multiple point patterns. The illumination pattern may include a periodic laser point grid. Illumination unit 118 may include at least one mask adapted to generate the illumination pattern from at least one beam of light generated by illumination unit 118.
[0188] Skin detection step 116 includes determining at least one second image, also referred to as a reflection image, using camera 112. The method may include determining multiple second images. Reflection features of the multiple second images may be used for skin detection in step b) and / or for 3D detection in step c). Reflection features may be features in an image plane generated by a scene in response to illumination (specifically having at least one illumination feature). Each reflection feature includes at least one beam profile, also referred to as a reflection beam profile. The beam profile of a reflection feature may generally refer to at least one intensity distribution of the reflection feature, such as the intensity distribution of a light spot on an optical sensor, as a function of pixels. The beam profile may be selected from the group consisting of: trapezoidal beam profiles; triangular beam profiles; conical beam profiles; and linear combinations of Gaussian beam profiles.
[0189] Evaluation of the second image may include identifying reflection features of the second image. Processing unit 114 may be configured to perform at least one image analysis and / or image processing to identify reflection features. Image analysis and / or image processing may use at least one feature detection algorithm. Image analysis and / or image processing may include one or more of the following: filtering; selecting at least one region of interest; forming a difference image between an image created by a sensor signal and at least one offset; inverting a sensor signal by inverting an image created by a sensor signal; forming a difference image between images created by a sensor signal at different times; background correction; decomposition into color channels; decomposition into hue, saturation, and luminance channels; frequency decomposition; singular value decomposition; applying a speckle detector; applying a corner detector; applying a determinant of a Hessian filter; applying a region detector based on principle curvature; applying a maximum stable extremum region detector; applying a generalized Hough transform; applying a ridge detector; applying an affine invariant feature detector; applying an affine adaptive interest point operator; applying a Harris affine region detector; applying a Hessian affine region detector. Detectors; application of scale-invariant feature transform; application of scale-space extremum detector; application of local feature detector; application of accelerated robust feature algorithm; application of gradient location and orientation histogram algorithm; application of histogram of orientation gradient descriptor; application of Deriche edge detector; application of differential edge detector; application of spatiotemporal interest point detector; application of Moravec corner detector; application of Canny edge detector; application of Gaussian Laplacian filter; application of Gaussian difference filter; application of Sobel operator; application of Laplacian operator; application of Scharr operator; application of Prewitt operator; application of Roberts operator; application of Kirsch operator; application of high-pass filter; application of low-pass filter; application of Fourier transform; application of Radon transform; application of Hough transform; application of wavelet transform; thresholding; creation of binary image. The region of interest can be determined manually by the user or automatically, for example, by identifying features within the image generated by optical sensor 126.
[0190] For example, illumination unit 118 can be configured to generate and / or project point clouds, such that multiple illuminated regions are generated on optical sensor 126 (e.g., a CMOS detector). Additionally, interference may exist on optical sensor 126, such as interference due to speckle and / or external light and / or multiple reflections. Processing unit 114 may be adapted to determine at least one region of interest, such as one or more pixels illuminated by a beam, for determining the longitudinal coordinates of corresponding reflective features, which will be described in more detail below. For example, processing unit 114 may be adapted to perform filtering methods, such as speckle analysis and / or edge filtering and / or object recognition methods.
[0191] Processing unit 114 can be configured to perform at least one image correction. Image correction may include at least one background subtraction. Processing unit 114 may be adapted to remove the influence of background light from the beam profile, for example, by imaging without further illumination.
[0192] Processing unit 114 can be configured to determine the beam profile of a corresponding reflection feature. Determining the beam profile may include identifying and / or selecting at least one reflection feature provided by optical sensor 126 and evaluating at least one intensity distribution of that reflection feature. As an example, the intensity distribution can be determined using a region of an evaluation matrix, such as a three-dimensional intensity distribution or a two-dimensional intensity distribution, for example, along an axis or line passing through the matrix. As an example, the illumination center of the beam can be determined, for example, by determining at least one pixel with the highest illumination, and a cross-sectional axis passing through the illumination center can be selected. The intensity distribution can be an intensity distribution that is a function of coordinates along a cross-sectional axis passing through the illumination center. Other evaluation algorithms are also feasible.
[0193] Processing unit 114 is configured to determine first beam profile information of at least one reflection feature by analyzing the beam profile of at least one reflection feature located within an image region of a second image corresponding to an image region of a first image including the identified geometric features. The method may include identifying an image region of the second image corresponding to an image region of the first image including the identified geometric features. Specifically, the method may include matching pixels of the first and second images and selecting pixels of the second image corresponding to the image region of the first image including the identified geometric features. The method may also include additional reflection features located outside the image region of the second image.
[0194] The beam profile information may be or may include any information and / or properties derived from and / or associated with the beam profile of the reflection features. The first and second beam profile information may be the same or different. For example, the first beam profile information may be intensity distribution, reflection profile, intensity center, or material features. For skin detection in step b) 116, beam profile analysis may be used. Specifically, beam profile analysis classifies materials using the reflection properties of coherent light projected onto the object surface. Material classification may be performed as described in WO 2020 / 187719, EP application 20159984.2 filed February 28, 2020, and / or EP application 20154961.5 filed January 31, 2020, the entire contents of which are incorporated herein by reference. Specifically, a periodic laser spot grid, such as the hexagonal grid described in EP application 20170905.2 filed April 22, 2020, is projected, and a reflection image is recorded using a camera. Feature-based methods can be used to analyze the beam profile of each reflective feature recorded by the camera. Refer to the description above for details on feature-based methods. Feature-based methods can be combined with machine learning methods, which allow for the parameterization of skin classification models. Alternatively or in combination, convolutional neural networks can be used to classify skin using reflective images as input.
[0195] Skin detection step 116 may include determining at least one material property of a reflective feature based on beam profile information using processing unit 114. Specifically, processing unit 114 is configured to identify reflective features that will be generated by irradiating biological tissue (particularly human skin) where the reflective beam profile meets at least one predetermined or predefined criterion. The at least one predetermined or predefined criterion may be at least one property and / or value suitable for distinguishing biological tissue (particularly human skin) from other materials. The predetermined or predefined criterion may be or may include at least one predetermined or predefined value and / or threshold and / or threshold range relating to material properties. If the reflective beam profile meets at least one predetermined or predefined criterion, the reflective feature may be indicated as being generated by biological tissue. The processing unit is configured to otherwise identify the reflective feature as non-skin. Specifically, processing unit 114 may be configured for skin detection, particularly for identifying whether a detected face is human skin. If the material is biological tissue, particularly human skin, the identification may include determining and / or verifying whether the surface to be examined or tested is or includes biological tissue (particularly human skin), and / or distinguishing biological tissue (particularly human skin) from other tissues (particularly other surfaces). The method according to the invention allows for the differentiation of human skin from one or more of inorganic tissues, metal surfaces, plastic surfaces, foams, paper, wood, displays, screens, and fabrics. The method according to the invention also allows for the differentiation of human biological tissues from the surfaces of artificial or inanimate objects.
[0196] Processing unit 114 can be configured to determine the material property m of the surface emitting the reflection feature by evaluating the beam profile of the reflection feature. The material property can refer to at least one arbitrary property of a material configured to characterize and / or identify and / or classify the material. For example, a material property can be a property selected from the group consisting of: roughness, depth of light penetration into the material, properties characterizing the material as a biological or non-biological material, reflectivity, specular reflectivity, diffuse reflectivity, surface properties, measurement of translucency, scattering, particularly backscattering behavior, etc. At least one material property can be a property selected from the group consisting of: scattering coefficient, translucency, transparency, deviation of Lambertian surface reflection, speckle, etc. Determining at least one material property can include assigning the material property to the detected face. Processing unit 114 can include at least one database comprising a list and / or table of predefined and / or pre-determined material properties, such as a lookup list or lookup table. The list and / or table of material properties can be determined and / or generated by performing at least one test measurement, for example, by performing a material test using a sample with known material properties. The list and / or table of material properties can be determined and / or generated at the manufacturer's location and / or by the user. Material properties can also be assigned to a material classifier, such as one or more of the following: material name, material group, such as biological or non-biological material, translucent or non-translucent material, metallic or non-metallic, skin or non-skin, fur or non-fur, carpet or non-carpet, reflective or non-reflective, specular or non-specular, foam or non-foam, hair or non-hair, roughness group, etc. Processing unit 114 may include at least one database comprising lists and / or tables that include material properties and associated material names and / or material groups.
[0197] While feature-based methods are accurate enough to distinguish skin from surface-scattering materials, differentiating skin from carefully selected attack materials (which also involve volume scattering) is more challenging. Step b) 116 may include the use of artificial intelligence, specifically convolutional neural networks. Using the reflection image as input to the convolutional neural network can generate a classification model with sufficient accuracy to distinguish skin from other volume-scattering materials. Since only physically valid information is passed to the network by selecting important regions in the reflection image, a compact training dataset may be required. Furthermore, a very compact network architecture can be generated.
[0198] Specifically, in skin detection step 116, at least one parameterized skin classification model can be used. The parameterized skin classification model can be configured to classify skin and other materials using a second image as input. The skin classification model can be parameterized using one or more of machine learning, deep learning, neural networks, or other forms of artificial intelligence. Machine learning can include methods for automatically building models, particularly parameterizing models, using artificial intelligence (AI). The skin classification model can include a classification model configured to distinguish human skin from other materials. The properties of the skin can be determined by applying an optimization algorithm based on at least one optimization objective on the skin classification model. Machine learning can be based on at least one neural network, particularly a convolutional neural network. The weights and / or topology of the neural network can be predetermined and / or predefined. Specifically, machine learning can be used to train the skin classification model. The skin classification model can include at least one machine learning architecture and model parameters. For example, the machine learning architecture can be or can include one or more of the following: linear regression, logistic regression, random forest, Naive Bayes classification, nearest neighbor, neural network, convolutional neural network, generative adversarial network, support vector machine, or gradient boosting algorithm, etc. As used herein, training can include the process of building a skin classification model, specifically determining and / or updating the parameters of the skin classification model. The skin classification model can be at least partially data-driven. For example, the skin classification model can be based on experimental data, such as data determined by illuminating multiple people and artificial objects like masks and recording reflection patterns. Training can, for example, include using at least one training dataset, wherein the training dataset comprises images of multiple people and artificial objects with known material properties, particularly a second image.
[0199] Skin detection step 116 may include using at least one 2D face and facial landmark detection algorithm configured to provide at least two locations of feature points of a face. For example, the locations could be the eye location, forehead, or cheek. The 2D face and facial landmark detection algorithm can provide the locations of feature points of a face, such as the eye location. Because there are subtle differences in reflections in different areas of the face (e.g., the forehead or cheek), a region-specific model can be trained. In skin detection step 116, at least one region-specific parametric skin classification model is preferably used. The skin classification model may include multiple region-specific parametric skin classification models, for example, for different regions, and / or region-specific data may be used to train the skin classification model, for example, by filtering the images used for training. For example, for training, two different regions may be used, such as the eye-cheek region below the nose, and particularly if sufficient reflective features cannot be identified in that region, the forehead region may be used. However, other regions are also possible.
[0200] If the material property corresponds to at least one property characteristic of skin, the detected face is characterized as skin. Processing unit 114 can be configured to identify reflective features generated by illumination of biological tissue, particularly skin, if its corresponding material property satisfies at least one predetermined or predefined criterion. If the material property indicates "human skin," the reflective feature can be identified as generated by human skin. If the material property is within at least one threshold and / or at least one range, the reflective feature can be identified as generated by human skin. At least one threshold and / or range can be stored in a table or lookup table and can be determined, for example, empirically, and can be stored as an example in at least one data storage device of the processing unit. Processing unit 114 is configured to otherwise identify the reflective feature as background. Therefore, processing unit 114 can be configured to assign a material property, such as whether skin is present or not, to each projection point.
[0201] The 3D detection step 120 can be performed after the skin detection step 116 and / or the face detection step 110. However, other embodiments are also possible, in which the 3D detection step 120 is performed before the skin detection step 116 and / or the face detection step 110.
[0202] 3D detection step 120 includes determining second beam profile information for at least four reflection features by analyzing beam profiles of at least four reflection features located within an image region of a second image corresponding to an image region of a first image including the identified geometric features. The second beam profile information may include the quotient Q of the region of the beam profile.
[0203] Beam profile analysis may include beam profile evaluation and may include at least one mathematical operation and / or at least one comparison and / or at least symmetry and / or at least one filtering and / or at least one normalization. For example, beam profile analysis may include at least one of histogram analysis steps, calculation of difference measurements, application of neural networks, and application of machine learning algorithms. Processing unit 114 may be configured to symmetry and / or normalize and / or filter the beam profile, particularly to remove noise or asymmetry from recordings at large angles, recording edges, etc. Processing unit 114 may filter the beam profile by removing high spatial frequencies (e.g., through spatial frequency analysis and / or median filtering, etc.). Summarization may be performed by averaging all intensities at the intensity center of the beam spot and at the same distance from the center. Processing unit 114 may be configured to normalize the beam profile to the maximum intensity, particularly taking into account intensity differences due to recording distances. Processing unit 114 may be configured to remove the effects of background light from the beam profile, for example, by imaging in the absence of illumination.
[0204] Processing unit 114 can be configured to determine at least one longitudinal coordinate z of a corresponding reflection feature by analyzing the beam profile of a reflection feature located within an image region of a second image corresponding to an image region of a first image including the identified geometric features. DPR Processing unit 114 can be configured to determine the longitudinal coordinate z of the reflection feature by using a technique known as photon depth ratio (also referred to as beam profile analysis). DPR For technical references on depth-to-photon ratio (DPR), see WO2018 / 091649A1, WO 2018 / 091638A1 and WO 2018 / 091640A1, the entire contents of which are incorporated herein by reference.
[0205] The longitudinal coordinates of the reflection feature can be the distance between the optical sensor 126 and the point in the scene emitting the corresponding illumination feature. Analysis of the beam profile of one of the reflection features can include determining at least one first region and at least one second region of the beam profile. The first region of the beam profile can be region A1 and the second region of the beam profile can be region A2. The processing unit 114 can be configured to integrate the first and second regions. The processing unit 114 can be configured to derive a combined signal, particularly a quotient Q, by one or more of the following: dividing the integrated first region and the integrated second region, dividing the integrated first region and the integrated second region by a multiple, or dividing the integrated first region and the integrated second region by a linear combination. The processing unit 114 can be configured to determine at least two regions of the beam profile and / or segment the beam profile into at least two segments comprising different regions of the beam profile, wherein overlap of regions is possible as long as the regions are not identical. For example, the processing unit 114 can be configured to determine multiple regions, such as two, three, four, five, or up to ten regions. The processing unit 114 can be configured to segment a light spot into at least two regions of the beam profile and / or segment the beam profile into at least two segments comprising different regions of the beam profile. Processing unit 114 can be configured to determine integrals of the beam profile over corresponding regions for at least two regions. The processing unit can be configured to compare at least two of the determined integrals. Specifically, processing unit 114 can be configured to determine at least one first region and at least one second region of the beam profile. The region of the beam profile can be any region of the beam profile at the location of the optical sensor used to determine the quotient Q. The first region and the second region of the beam profile can be one or both of adjacent or overlapping regions. The first region and the second region of the beam profile can be different in area. For example, processing unit 114 can be configured to divide the sensor region of a CMOS sensor into at least two sub-regions, wherein the processing unit can be configured to divide the sensor region of the CMOS sensor into at least one left portion and at least one right portion and / or at least one upper portion and at least one lower portion and / or at least one inner portion and at least one outer portion. Alternatively, camera 112 may include at least two optical sensors 126, wherein the photosensitive regions of the first optical sensor 126 and the second optical sensor 126 may be arranged such that the first optical sensor 126 is adapted to determine a first region of the beam profile of the reflection characteristics, and the second optical sensor 126 is adapted to determine a second region of the beam profile of the reflection characteristics. Processing unit 114 may be adapted to integrate the first and second regions. Processing unit 114 may be configured to determine the longitudinal coordinate using at least one predetermined relationship between the quotient Q and the longitudinal coordinate. The predetermined relationship may be one or more of empirical relationships, semi-empirical relationships, and analytically derived relationships.The processing unit 114 may include at least one data storage device for storing predetermined relationships, such as lookup lists or lookup tables.
[0206] The 3D detection step may include determining at least one depth level from second beam profile information of the reflected feature using a processing unit.
[0207] Processing unit 114 can be configured to determine a depth map of at least a portion of a scene by determining at least one depth information of reflective features located within an image region of a second image corresponding to an image region of a first image including the identified geometric features. Processing unit 114 can be configured to determine the depth information of the reflective features using one or more of the following techniques: photon depth ratio, structured light, beam profile analysis, time-of-flight, shape from motion, depth from focus, triangulation, depth from defocus, and stereo sensor. The depth map can be a sparsely filled depth map including a few entries. Alternatively, the depth map may be crowded, including a large number of entries.
[0208] If the depth level deviates from a predetermined or predefined depth level of the planar object, the detected face is characterized as a 3D object. Step c) 120 may include using 3D topological data of the face in front of the camera. The method may include determining curvature based on at least four reflection features located within an image region of a second image corresponding to an image region of a first image including the identified geometric features. The method may include comparing the curvature determined from the at least four reflection features with a predetermined or predefined depth level of the planar object. If the curvature exceeds the assumed curvature of the planar object, the detected face may be characterized as a 3D object; otherwise, it is characterized as a planar object. The predetermined or predefined depth level of the planar object may be stored in at least one data memory of the processing unit, such as a lookup list or lookup table. The predetermined or predefined depth level of the planar object may be determined experimentally and / or may be a theoretical level of the planar object. The predetermined or predefined depth level of the planar object may be at least one limit of at least one curvature and / or a range of at least one curvature.
[0209] The 3D features identified in step c) 120 allow for the differentiation between high-quality photographs and 3D facial structures. The combination of steps b) 116 and c) 120 allows for enhanced reliability of authentication against attacks. 3D features can be combined with material features to improve security levels. Since the same computational pipeline can be used to generate the input data for skin classification and the 3D point cloud, both properties can be computed from the same frame with lower computational cost.
[0210] Preferably, authentication step 122 can be performed after steps a) 110, b) 116, and c) 120. Authentication step 122 can be performed partially after each of steps a) through c). Authentication can be aborted if no face is detected in step a) 110 and / or if it is determined in step b) 116 that the reflection features are not generated by skin and / or if the depth map involves a planar object in step c) 120. The authentication step includes authenticating the detected face by using at least one authentication unit if the face detected in step b) 116 is characterized as skin and the face detected in step c) 122 is characterized as a 3D object.
[0211] Steps a) through d) can be performed using at least one device, such as at least one mobile device 124, like a mobile phone, smartphone, etc., where access to the device is protected by using facial authentication. Other devices are also possible, such as access control devices that control access to buildings, machines, cars, etc. The method may include allowing access to the device if the detected face is authenticated.
[0212] The method may include at least one registration step. In the registration step, a user of the device may be registered. Registration may include a process of registering and / or signing up and / or instructing the user to use the device subsequently. Typically, registration may be performed upon first use of the device and / or startup of the device. However, embodiments in which multiple users may be registered consecutively, making it feasible to perform and / or repeat registration at any time during device use, are possible. Registration may include generating a user account and / or user profile. Registration may include inputting and storing user data, particularly image data, via at least one user interface. Specifically, at least one 2D image of the user is stored in at least one database. The registration step may include imaging at least one image of the user, particularly multiple images. Images may be recorded from different orientations and / or the user may change their orientation. Additionally, the registration step may include generating at least one 3D image and / or depth map for the user, which may be used for comparison in step d). The database may be the device's database, such as the database of processing unit 114, and / or may be an external database such as the cloud. The method includes identifying the user by comparing the user's 2D image with a first image. The method according to the invention can allow for a significant improvement in the presentation attack detection capability of biometric authentication methods. To improve overall authentication, in addition to the user's 2D image, personal material fingerprints and 3D topological features can be stored during the registration process. This allows for multi-factor authentication within a single device using 2D, 3D, and material-derived features.
[0213] The method using beam profile analysis technology according to the present invention provides a concept for reliably detecting human skin by analyzing the reflection of laser spots on the face (particularly in the NIR range) and distinguishing it from reflections from attack materials produced by mimicking the face. Furthermore, beam profile analysis simultaneously provides depth information by analyzing the same camera frame. Therefore, 3D and skin-safe features can be provided using the exact same technology.
[0214] Since 2D images of the face can also be recorded simply by turning off the laser illumination, a completely secure facial recognition pipeline can be established to solve the above problems.
[0215] As the laser wavelength shifts into the NIR region, the reflectivity of human skin of different skin tones becomes more similar. The differences are minimal at a wavelength of 940 nm. Therefore, skin color does not contribute to skin authentication.
[0216] It may not be necessary to perform time-consuming analysis on a series of frames, since the presentation attack detection (via skin classification) is provided by only one frame. The time frame used to perform the full method can be ≤500 ms, preferably ≤250 ms. However, embodiments in which multiple frames can be used to perform skin detection may be feasible. Depending on the confidence level of identifying reflection features in the second image and the speed of the method, the method may include sampling reflection features on multiple frames to achieve a more stable classification.
[0217] Figure 3 The experimental results are shown, particularly the density as a function of skin score. The x-axis displays the score, and the y-axis displays the frequency. The score is a measure of classification quality, ranging from 0 to 1, where 1 indicates very high skin similarity and 0 indicates very low skin similarity. The decision threshold is likely around 0.5. A reference distribution of the skin scores for real presentations was generated using 10 subjects. Skin scores were also recorded for presentation attacks (PAs) at levels A, B, and C (as defined in relevant ISO standards). The experimental setup (the evaluation target, ToE) included proprietary hardware devices, such as... Figure 2 As shown, this includes the necessary sensors and the computing platform to execute the PAD software. ToE was tested using six Class A PAIs (Presentation Attack Tools), five Class B PAIs, and one Class C PAI. For each PAI category, ten PAIs were used. The types of PAIs used in this study are listed in the table below. In the table, APCER is the attack presentation classification error rate, which is the number of successful attacks divided by the total number of attacks. 100. In the table, BPCER represents the classification error rate for the true (BonaFide) state, which is the number of rejected unlock attempts divided by the total number of unlock attempts. Attacks of type 100, A, and B were based on 2D PAI, while the type C attack was based on a 3D mask. For the type C attack, a custom-made rigid mask constructed using a 3D printer was used. A test subject group of 10 subjects was used to obtain a reference distribution of the realistically presented skin score.
[0218]
[0219] Experiments using these PAIs demonstrate that the two types of presentation (real or PA) can be clearly distinguished based on skin scores. The method according to the invention can clearly differentiate between paper, 3D printing, and skin.
[0220] Reference Number List
[0221] 110 Facial Detection Steps
[0222] 112 cameras
[0223] 114 Processing Units
[0224] 116 Skin Detection Steps
[0225] 118 lighting units
[0226] 120 3D Inspection Steps
[0227] 122 Authentication Steps
[0228] 124 Mobile Devices
[0229] 126 Optical Sensors
Claims
1. A method for facial authentication, comprising the following steps: a) At least one face detection step (110), wherein the face detection step (110) includes: using at least one camera (112) to determine at least one first image, wherein the first image includes at least one two-dimensional image of a scene suspected of including a face, wherein the face detection step (110) further includes: using at least one processing unit (114) to detect the face in the first image by identifying at least one predetermined geometric feature characteristic of the face in the first image; b) At least one skin detection step (116), wherein the skin detection step (116) includes: using at least one illumination unit (118) to project at least one illumination pattern including a plurality of illumination features onto the scene, and using the at least one camera (112) to determine at least one second image, wherein the second image includes a plurality of reflection features generated by the scene in response to illumination of the illumination features, wherein each of the reflection features includes at least one beam profile, wherein the skin detection step further includes: using the processing unit (114) to determine first beam profile information of the at least one reflection feature by analyzing the beam profile of at least one reflection feature located in an image region of the second image corresponding to an image region of the first image including the identified geometric features, and determining at least one material property of the reflection feature from the first beam profile information, wherein if the material property corresponds to at least one property characteristic for skin, the detected face is characterized as skin; c) At least one 3D detection step (120), wherein the 3D detection step (120) includes: using the processing unit (114) to determine second beam profile information of at least four reflection features by analyzing beam profiles of at least four reflection features located in the image region of the second image corresponding to the image region of the first image including the identified geometric features, and determining at least one depth level from the second beam profile information of the reflection features, wherein if the depth level deviates from a predetermined depth level of a planar object, the detected face is characterized as a 3D object; d) At least one authentication step (122), wherein the authentication step (122) includes: authenticating the detected face by using at least one authentication unit if the face detected in at least one skin detection step (116) in step b) is characterized as skin and the face detected in at least one 3D detection step (120) in step c) is characterized as a 3D object.
2. The method according to claim 1, wherein, Steps a) to d) are performed using at least one device, wherein access to the device is protected by using facial authentication, wherein the method includes: allowing access to the device if the detected face is authenticated.
3. The method according to claim 2, wherein, The method includes at least one registration step, wherein, in the registration step, a user of the device is registered, wherein at least one 2D image of the user is stored in at least one database, and wherein the method includes: identifying the user by comparing the 2D image of the user with a first image.
4. The method according to claim 3, wherein, In the skin detection step, at least one parametric skin classification model is used, wherein the parametric skin classification model is configured to classify skin and other materials by using the second image as input.
5. The method according to claim 4, wherein, The skin classification model is parameterized by using machine learning, wherein the properties of the skin are determined by applying an optimization algorithm on the skin classification model in terms of at least one optimization objective.
6. The method according to claim 4 or 5, wherein, The skin detection step includes: using at least one 2D face and facial landmark detection algorithm, the 2D face and facial landmark detection algorithm being configured to provide at least two locations of feature points of a face, wherein, in the skin detection step (116), at least one region-specific parametric skin classification model is used.
7. The method according to claim 1, wherein, The lighting pattern comprises a periodic grid of laser spots.
8. The method according to claim 1, wherein, The illumination feature has a wavelength in the near-infrared (NIR) range.
9. The method according to claim 8, wherein, The illumination feature has a wavelength of 940 nm.
10. The method according to claim 1, wherein, A plurality of second images are determined, wherein the reflection features of the plurality of second images are used for skin detection in at least one skin detection step (116) in step b) and / or for 3D detection in at least one 3D detection step (120) in step c).
11. The method according to claim 1, wherein, The camera (112) includes at least one near-infrared camera.
12. A computer program for facial authentication, configured to, when executed on a computer or computer network, cause the computer or computer network to perform, fully or partially, the method according to any one of claims 1 to 11, wherein, The computer program is configured to perform and / or implement at least steps a) to d) of the method according to any one of claims 1 to 11.
13. A computer-readable storage medium comprising instructions that, when executed by a computer or computer network, cause at least the following steps (a) to (d) of the method according to any one of claims 1 to 11 relating to the method to a method.
14. A non-transitory computer-readable medium comprising instructions that, when executed by one or more processors, cause the one or more processors to perform at least steps a) to d) of the method according to any one of claims 1 to 11 relating to the method.
15. A mobile device (124) comprising at least one camera (112), at least one illumination unit (118), and at least one processing unit (114), said mobile device (124) being configured to perform at least steps a) to c) of a method for facial authentication according to any one of claims 1 to 11 relating to the method, or steps a) to d) of the method.
16. The use of the method according to any one of claims 1 to 11 for detecting biological presentation attacks.
Citation Information
Patent Citations
Facial authentication systems and methods utilizing time of flight sensing
US20190213309A1
Detector for optically detecting at least one object
WO2018091638A1
Detector for optically detecting at least one object
WO2018091640A2
Detector for optically detecting at least one object
WO2018091649A1
Detector for identifying at least one material property
WO2020187719A1