In vivo determination of domain transformation based on thermal images
By combining a thermal imaging camera with domain transformation and machine learning models, the problem of detecting complex spoofing tools in existing technologies has been solved, enabling effective identification of facial spoofing at long distances and on low-cost devices, and is suitable for passive facial recognition systems.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- THALES DIS FRANCE SA
- Filing Date
- 2024-07-12
- Publication Date
- 2026-05-01
AI Technical Summary
Existing methods for detecting facial presentation attacks are ineffective at detecting sophisticated deception tools such as silicone or latex full-face masks. Furthermore, traditional methods are not effective at long distances or on low-cost devices and cannot perform passive detection without the need for the target to actively cooperate.
After receiving images from a thermal imaging camera, performing face detection and key point recognition, the images are transformed from the thermal domain to another spectral domain, such as the visible domain, using domain transformation technology. Liveness determination is then performed using machine learning models, including generative adversarial networks and super-resolution processing, to generate realistic visible images for classification.
It achieves robust detection of various PAIs, effectively identifies facial spoofing at long distances and on low-cost devices, improves detection accuracy and adaptability, and is suitable for passive facial recognition systems.
Smart Images

Figure CN121970097A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of facial processing, and more specifically to facial presentation attack detection, which aims to determine whether a face is real / live, originating from a target person at the capture point, or is deceptive / forged. Background Technology
[0002] More and more applications and services now rely on facial recognition (FR) as a pre-processing step to ensure that users are authorized to use resources such as services / applications.
[0003] This is the case, for example, in the telecommunications sector (e.g., when unlocking a phone), when accessing a physical location (e.g., in an airport), in biometric systems used for authentication, or in digital identification (e.g., for banking services).
[0004] Among the services mentioned above, facial recognition is becoming one of the preferred biometric modalities.
[0005] Once a user's face is captured, image frames are typically processed and compared with stored images tagged with the user's identity in order to infer the user's identity.
[0006] In order to access services / applications featuring real people, attackers may carry out presentation attacks, referred to below as PA. PA may include presenting a spoofed face to an identification system.
[0007] To address these presentation attack (PA) attacks, presentation attack detection methods have been developed and applied before facial recognition to ensure that the identified real person is genuine and alive, and not a deception.
[0008] Some attack detection methods or systems are called proactive, in which they require the target person to perform a series of one or more instructions, thereby requesting some type of cooperative behavior, such as blinking, mouth / lip movement and / or head turning.
[0009] However, the drawback of proactive methods / systems is their intrusiveness, which is a problem in security applications where they need to be identified without the target being aware of them.
[0010] In contrast, passive presentation attack detection does not require any cooperation from the target person, is therefore non-intrusive, and processes images or frames acquired by capturing devices such as cameras.
[0011] Modern methods of facial presentation attacks (PAs) are causing problems for rendering systems (FRs), thus posing a major vulnerability risk to services / applications that rely on these FRs. In fact, modern methods of facial presentation attacks benefit from the widespread availability and high quality of tools, referred to below as PAIs (Presentation Attack Tools), which include:
[0012] - Printed or cropped images;
[0013] - More complex PAIs, such as video playback or folded paper masks;
[0014] - Expensive and advanced PAIs, such as silicone or latex full-face masks.
[0015] Current presentation attack detection methods are adapted to the sensors available in every FR system. Most presentation attack detection methods must rely solely on standard red-green-blue (RGB) cameras or near-infrared (NIR) cameras.
[0016] It should be noted that the infrared spectrum includes four spectral bands: NIR, "short-wave infrared" (SWIR), "mid-wave infrared" (MWIR), and "long-wave infrared" (LWIR). MWIR and LWIR are also referred to as "thermal".
[0017] However, presentation attack detection methods based on RGB cameras are inefficient at detecting presentation attacks based on elaborate video replay attacks.
[0018] Other methods implemented in FR systems include time-of-flight cameras or structured light. These sensors have the benefit of adding three-dimensional (3D) information for analysis, making the detection of PAIs, such as images, much easier.
[0019] However, these sensors are very expensive, and they lose their effectiveness when the target is at a distance. They are also difficult to integrate with passive presentation attack detection methods.
[0020] The publication “A Study on Presentation Attack Detection in Thermal Infrared” (Kowalski, Military University of Technology, Sensors 20(14):3988, July 2020) proposes a method for facial temperature measurement and visualization for PA detection. This relies on a two-step deep learning-based approach, consisting of a head detector followed by a deep learning classifier.
[0021] However, this method encounters limitations when it comes to more complex PAs, such as full-face PAs based on silicone or latex.
[0022] For example, refer to Figure 1 Three images are shown of a person wearing a full-face latex mask as a PAI (Presentation Attack Impersonator). The first image, 100.1, is a thermal image acquired by a thermal infrared camera, such as those used in the prior art publications mentioned above. In the first image 100.1, the target person begins wearing the mask before the first image 100.1 is acquired, meaning the mask is cold and the resulting thermal spectral contrast is low: a thermally uniform shape and temperature are obtained, which appears to be a forgery. Therefore, the first image 100.1 can be identified as a forgery by thermal image-based presentation attack detection methods, such as those described in the referenced publications above.
[0023] However, in the second image 100.2, the same target person is shown wearing a full-face latex mask but with a heated mask, which was worn by the target person before the second image 100.2 was captured by the thermal infrared camera.
[0024] The third image 100.3 shows an image of the same person captured by a camera that captures the image in the visible field.
[0025] As can be observed, the second image 100.2 includes more temperature variations and lower thermal uniformity. Therefore, when processing the second image 100.2, PA detection methods using thermal imaging cameras (such as those described in the publications above) may fail to detect that the target person's face is being deceived.
[0026] Therefore, there is a need for a presentation attack detection method capable of detecting several types of PAIs, including the most complex types such as full-face latex or silicone masks. Preferably, the presentation attack detection method can be passive and can be applied to targets at a distance.
[0027] The present invention aims to improve the current situation. Summary of the Invention
[0028] This invention provides solutions to the aforementioned problems through the presentation attack detection method according to claim 1, the computer program according to claim 15, and the presentation attack detection device according to claim 16. Preferred embodiments of the invention are defined in the dependent claims.
[0029] A first aspect of the present invention relates to a presentation attack detection method, the method comprising:
[0030] - Receive at least one thermal image from a thermal imaging camera;
[0031] - Determine whether the target person's face is detected in the thermal image;
[0032] -If the face of the target person is detected in the thermal image, a domain transformation is performed based on the thermal image to obtain an image in a second spectral domain, which is different from the thermal domain;
[0033] -Based on the image in the second spectral domain, determine the first category indicating whether the target person is alive or deceiving;
[0034] -A final decision indicating the liveness of the target person is made based on the determined first category and / or based on whether the target person's face is detected in the thermal image.
[0035] This invention integrates continuous face detection and domain transformation, which allows for the detection of a wide variety of PAIs, including full-face masks made of latex or silicone, due to the detection of thermal emission interference. In fact, the image synthesized in the second spectral domain highlights relevant information revealing / denying the liveness of the target person. Because this invention relies on several decisions aggregated into a final decision, the method is robust to various PAIs.
[0036] According to some implementations, the method may further include performing canonical face alignment of the thermal image to obtain a first cropped thermal image, and the domain transformation may be performed on the first cropped thermal image to obtain the image in the second spectral domain.
[0037] Therefore, the target person does not need to be positioned at a given distance and in front of the thermal imaging camera, and the present invention can be integrated into a passive facial recognition system.
[0038] As an additional step, the method may further include detecting at least one key point of the target person's face if the face of the target person is detected in the thermal image, the domain transformation may be performed if the face of the target person is detected and if at least one key point of the target person's face is detected, and the final decision may be further based on whether at least one key point of the target person's face is detected in the thermal image.
[0039] Therefore, the detection of key points can be taken into account together with facial detection to detect some PAIs, which makes the present invention more robust for the detection of a wide variety of PAIs.
[0040] As a supplement, this specification allows for face alignment based on at least one detected key point on the target person's face. This enables improved accuracy of face alignment and avoids requiring the target person to be precisely positioned in front of the thermal imaging camera.
[0041] According to some implementation schemes, the domain transformation can be performed by applying a model configured to determine the image in the second spectral domain in response to the corresponding thermal image received as input.
[0042] This model can be obtained through machine learning, which allows for improved accuracy associated with domain transformation operations. The model can be, for example, a generative model. Alternatively, the model can be a diffusion model.
[0043] As a supplement, the model can be configured to perform super-resolution to receive a thermal image with a first resolution as input, and to output an image with a second resolution in the second spectral domain that is higher than the first resolution.
[0044] This enables the processing of low-resolution thermal images, making the invention compatible with low-cost thermal imaging cameras and / or allowing target personnel to be positioned away from the thermal imaging camera. Performing domain transformation at different resolutions is synonymous with real-world deployment because it does not require personnel to stand still at a specific distance, for example.
[0045] Alternatively or as a supplement, the model can have an encoder-decoder structure based on a pyramidal architecture. This allows for multi-scale analysis to generate high-resolution images in a second spectral domain. This then improves the accuracy associated with the determination of the first category.
[0046] Alternatively or additionally, the method may also include a preliminary step of training the model using a generative adversarial network, and the generative model may be trained by a discriminator using at least one adversarial loss function. This enables the generation of realistic images in a second spectral domain.
[0047] As a supplement, the model can be trained by the discriminator using a global adversarial loss function and at least one local adversarial loss function. This enables the generation of globally realistic images in a second spectral domain, as well as realistic images for specific parts of the face, such as biometrically significant parts.
[0048] Alternatively or additionally, the model can be further trained based on at least one of the following loss functions: L1 loss function, perceptual loss function, and identity loss function. This enables the perceptual rendering of the image to be preserved, biometric identity-specific features to be retained, and consistent age-gender reconstruction to be implemented.
[0049] According to some implementation schemes, the model may be trained by supervised learning on a training dataset that includes a plurality of thermal images with a first resolution associated with an image with a second resolution in the second spectral domain, the second resolution being higher than the first resolution.
[0050] Therefore, this model can process thermal images as inputs with variable low resolution, allowing for the determination of the liveness of a target person located at a variable distance from the thermal imaging camera. Thus, the method is robust and compatible with low-cost sensors. It is adaptable to real-world scenarios because face detection can operate under a wide range of conditions, and the domain transformation accepts various thermal image resolutions / qualities.
[0051] According to some implementations, this second spectral domain can be the visible domain. Therefore, the synthesized visible image brings relevant information that can be discerned by classification models used for PADs, such as deep neural network classifiers.
[0052] According to some implementations, the method may further include, if the target person's face is detected in the thermal image, determining a second category based on the thermal image indicating whether the target person is alive or deceiving; and the final decision may be further based on the second category. This enables the detection of a wide variety of PAIs. Furthermore, it enables the differentiation of several types of PAIs when the target person's face is detected as deceiving.
[0053] As a supplement, based on the corresponding classification model, the first category can be determined by the first classification module, and the second category can be determined by the second classification module.
[0054] This enables the detection of a wide variety of PAIs based on different classification models that process images in different spectral domains. Therefore, this method is robust to very different types of PAIs.
[0055] A second aspect of the invention relates to a computer program comprising instructions arranged to implement the method according to the first aspect of the invention when executed by a processor.
[0056] A third aspect of the present invention relates to a presentation attack detection device, the presentation attack detection device comprising:
[0057] - A receiving interface for receiving at least one thermal image;
[0058] - A face detection module configured to determine whether the face of a target person is detected in the thermal image;
[0059] - A domain conversion module, which is configured to perform a domain conversion based on the thermal image to obtain an image in a second spectral domain, which is different from the thermal domain;
[0060] - A first classification module, which is configured to determine a first category indicating whether the target person is alive or deceiving based on the image in the second spectral domain;
[0061] - A decision module configured to determine a final decision indicating the liveness of the target person based on a determined first category and / or based on whether the target person's face is detected in the thermal image.
[0062] All features and / or steps of the methods described herein (including the claims, description, and drawings) may be combined in any combination except for combinations of such mutually exclusive features and / or steps. Attached Figure Description
[0063] Referring to the accompanying drawings, these and other features and advantages of the invention will become clear from the detailed description of the invention, which will become apparent from the preferred embodiments of the invention, which are given by way of example only and are not limited thereto.
[0064] Figure 1 The image shows both thermal and visible images of a target wearing a presentation attack tool, such as a full-face mask.
[0065] Figure 2 This figure illustrates a presentation attack detection system according to some embodiments of the present invention.
[0066] Figure 3 This figure illustrates a domain translation module according to some embodiments of the present invention.
[0067] Figure 4 The figure illustrates the structure of a discriminator for a generative adversarial network used to train a generative model of a domain transformation module according to some embodiments of the present invention.
[0068] Figure 5 This figure shows a thermal image of a rendering attack based on a two-dimensional tool.
[0069] Figure 6 This image shows a thermal image of a mask-based presentation attack.
[0070] Figure 7 This figure is a diagram illustrating the steps of a presentation attack detection method according to some embodiments of the present invention. Detailed Implementation
[0071] As will be understood by those skilled in the art, aspects of the present invention may be embodied as presenting attack detection methods, apparatus or computer programs.
[0072] Figure 1 A system 20 for attack detection is shown according to some embodiments of the present invention.
[0073] The system includes a frame capture device 200 (such as a camera or video camera) arranged to acquire at least one thermal image 201, and a presentation attack detection device 210. The frame capture device 200 is a thermal imaging camera 200 according to the present invention.
[0074] As previously explained, thermal image 201 is an image in the MWIR or LWIR spectral domain.
[0075] The thermal imaging camera 200 can acquire a series of thermal images 201 at a given frequency. For example, when the thermal images 201 are video frames, the frequency can be higher than one frame per second (fps), such as 24 frames per second. Alternatively, the thermal imaging camera 200 is arranged to acquire thermal images 201 at a fixed frequency of less than 1 fps, or is arranged to acquire thermal images 201 when triggered by an event (e.g., when movement is detected).
[0076] The thermal imaging camera 200 is arranged to acquire thermal images representing a scene within the camera's field of view, which may be static or dynamic. There is no limitation on the resolution of the thermal image 201. The resolution depends on the sensitivity and format of the sensor in the thermal imaging camera 200.
[0077] The thermal imaging camera 200 can be installed at an end-user site, which can be at a building entrance, in front of a door, at a security gate in an airport, or at any other physical location.
[0078] The scene represented in at least some frames of the sequence represents at least the face of the target person. As will be understood below, some embodiments of the invention allow the presentation attack detection device 210 to process thermal images 201 for a wide range of distances between the target person and the thermal imaging camera 200, rather than just for a fixed distance where the target person must be precisely positioned in front of the thermal imaging camera 200. This allows the presentation attack detection system 20 to be used in passive applications that do not require active cooperation from the target person.
[0079] The attack detection device 210 may include an input interface 230 arranged to receive a thermal image 201 from a camera 200 or a series of thermal images 201 from the camera.
[0080] There are no restrictions on the presentation attack detection device 210, which can be any device including processing power or processor and non-transitory memory.
[0081] The presentation attack detection device 210 may be incorporated, for example, into a server, desktop computer, smart tablet computer, smartphone, mobile internet device, personal digital assistant, wearable device, image capture device, or any combination thereof. The presentation attack detection device 210 includes fixed-function hardware logic, configurable logic, logic instructions, etc., or any combination thereof.
[0082] Furthermore, there are no restrictions on the communication link between the thermal imaging camera 200 and the attack detection device 210, which transmits thermal images thereon. The communication link can be, for example, a wireless link or a wired link. For example, wired protocols may include RS-232, RS-422, RS-485, I2C, SPI, IEEE 802.3, and TCP / IP. Wireless protocols may include IEEE 802.11a / b / g / n, Bluetooth, Bluetooth Low Energy (BLE), FeliCa, Zigbee, GSM, LTE, 3G, 4G, 5G, RFID, and NFC.
[0083] The attack detection device 210 includes multiple modules 211 to 218, some of which are optional as explained below. These modules are either hardware modules or software modules.
[0084] The face detection module 211 is arranged to determine whether the face of the target person is included in the thermal image 201 received from the thermal imaging camera.
[0085] There are no restrictions on the face detection module 211, which may be based on a first model arranged to classify input images into two categories: "face" and "no face." The first model may be constructed analytically or based on machine learning. The first model may be, for example, a machine learning model, such as a neural network, trained through supervised learning on a first set of training data including first reference thermal images. Each first reference thermal image is associated with the label "face" if it includes the face of the target person, and with the label "no face" if it does not include any face of the target person.
[0086] The first model can be stored in the memory 220 of the attack detection device 210 or in the internal memory of the face detection module 211.
[0087] Preferably, the first model is trained to be robust to several conditions such as the target person's pose, expression, occlusion, poor thermal image quality, and long-range distance. For this purpose, the first set of training data includes reference thermal images acquired under different conditions of the target person's pose, expression, occlusion, quality, and distance.
[0088] Therefore, the face detection module 211 is arranged to receive the thermal image and to determine whether the thermal image includes a face. If not, the thermal image can be discarded, and the decision module 216 can be indicated with the category "no face". If the face of the target person is detected, the face detection module 211 sends the thermal image 201 to the next module for further processing. Furthermore, the face detection module 211 can send the category "face" to the decision module 216.
[0089] The attack detection device 210 may also include a facial key point detection module 212, which is arranged to determine at least one key point of the face detected by the facial detection module 211 in the thermal image 201, and preferably several key points.
[0090] Facial landmark detection typically involves at least one of the following landmarks, or any combination of the following landmarks: eyes, nose, mouth, and chin. In the following text, for illustrative purposes, facial landmark detection is considered to aim at identifying the landmarks listed above.
[0091] To this end, the face detection module 211 can process the thermal image 201 based on a second model, which is arranged to identify the aforementioned key points based on the thermal image 201 received from the face detection module 211. The second model can be constructed in an analytical manner or obtained based on machine learning.
[0092] Some examples of machine learning-based facial keypoint detection models are known, namely, predictive models in the form of neural networks, specifically convolutional neural networks (CNNs). A second model can be obtained by training a similar model using supervised learning based on a second set of training data, including thermal images labeled with reference coordinates of the keypoints.
[0093] The facial key point detection module 212 can also determine information indicating whether key points have been detected in the thermal image. If no key points are detected, the thermal image can be discarded, and the decision module 216 can be informed that "no key points were detected". If one or more key points have been detected, the facial key point detection module 212 sends the thermal image 201 to the next module for further processing. Furthermore, the facial key point detection module 212 can send the information "key points" to the decision module 216.
[0094] Therefore, both the face detection module 211 and the facial landmark detection module 212 can act as pre-filtering modules to detect some simple PAIs that do not require further processing.
[0095] According to some implementation schemes, the face detection module 211 and the facial key point detection module 212 are one and the same module, which can determine the key points of the face simultaneously when the face is detected in the thermal image 201, or can determine that the face of the target person is not included in the thermal image 201 based on a single face and key point detection model.
[0096] Once key points are detected in the thermal image, the key point detection module 212 sends the thermal image along with the coordinates of the key points (and optionally the geometric information of the key points) to the face alignment module 213. According to some embodiments of the invention, the thermal image and key points are also sent to an optional cropping module 217, or directly to a second classification module 218, which will be described below.
[0097] If some key points in the key points cannot be identified by the key point detection module 212, the attack detection device 210 may consider the target person's face to be partially or completely obscured and may discard the thermal image 201.
[0098] The face alignment module 213 is arranged to crop the thermal image 201 based on key points to perform canonical face alignment and obtain a first cropped thermal image to be sent to the domain conversion module 214.
[0099] Depending on the distance between the target person and the thermal imaging camera 200, the size of the target person's face in the thermal image 201 can vary. Therefore, when cropping the thermal image to obtain a first cropped thermal image aligned with the target person's standard face, the resolution of the first cropped thermal image depends on the size of the target person's face in the thermal image 201. The farther away the target person is, the lower the resolution of the first cropped thermal image.
[0100] Alternatively, the presentation attack detection device 210 is configured only to process thermal images 201 for which the distance between the thermal imaging camera 200 and the face of the target person is in a decreasing range, such as between a first distance of 1 meter and a second distance of 1.5 meters. If the target person is located at a distance less than 1 meter or greater than 1.5 meters, the thermal image 201 is not processed by the presentation attack detection device 210. According to this alternative, the resolution of the first cropped thermal image obtained by the face alignment module 213 is fixed.
[0101] Then, the first cropped thermal image is passed to the domain conversion module 214.
[0102] The cropping function performed by the face alignment module 213 can be performed by the key point detection module 212. In this case, the attack detection device 210 does not include the face alignment module 213, and the first cropped thermal image is transferred from the key point detection module 212 to the domain transformation module 214.
[0103] Domain transformation module 214 is configured to transform a cropped thermal image in a first spectral domain, such as MWIR or LWIR, into an image in a second spectral domain. For this purpose, the domain transformation may be based on a third model (which is a generative model), configured to receive the image in the first spectral domain and predict or generate a corresponding image in the second spectral domain. The third model may be a generative or diffusion model obtained through machine learning (e.g., through supervised learning based on a third set of data including multiple pairs of thermal images labeled with corresponding reference images in the second spectral domain). The third model is trained to minimize at least one loss function computed based on the predicted image in the second spectral domain and the reference image in the second spectral domain.
[0104] Preferably, the third model is arranged to process the first cropped thermal image with different resolutions, rather than just having a given fixed resolution.
[0105] Preferably, the second spectral domain can be the visible domain. However, the domain conversion module 214 is arranged to predict the image in a second spectral domain that is different from the visible domain. For example, the second spectral domain can be another infrared domain, such as NIR or SWIR.
[0106] Domain transformation enables the detection of any artifacts or defects that are not perceptible in the thermal domain.
[0107] In the following text, for illustrative purposes only, the second spectral domain is considered to be the visible domain.
[0108] Preferably, the third model is trained to retain the identity information contained in the first cropped thermal image in the predicted visible image.
[0109] Furthermore, according to some implementations, a third model is trained to enhance the resolution of the first cropped thermal image to obtain a high-resolution visible image, which may be low-resolution due to the thermal imaging camera 200 and the distance to the target person. The ability to predict a high-resolution image based on a low-resolution image is called super-resolution and will be explained below.
[0110] Reference Figure 3 A detailed example of a third model that allows both spectral domain transformation and super-resolution is further described.
[0111] The predicted visible image is passed from the domain transformation module 214 to the first classification module 215.
[0112] The first classification module 215 is configured to receive a visible image including the face of the target person as input, and is configured to classify the visible image into one of two categories, "live," if the target person is considered to be a real person, or to classify the visible image into one of two categories, "forgery," if PA is detected.
[0113] To this end, the first classification module 215 can implement a fourth model, which is configured to receive a visible image as input and to predict one of the two categories mentioned above, "deception" and "liveness". There are no restrictions on the fourth model; it can be analytical or it can be obtained based on machine learning.
[0114] For example, the fourth model obtained based on machine learning could be a neural network configured to extract features. The output layer of the fourth layer is configured to predict the category of liveness or deception based on features extracted from the visible image.
[0115] The fourth model can be trained through supervised learning based on a fourth set of training data, which includes reference images in the second spectral domain labeled with the categories of liveness or deception.
[0116] The first classification module 215 is configured to send the categories output by the fourth model to the decision module 216.
[0117] The decision module 216 is configured to determine the final decision of the target person of the thermal image 201 as either a live person or a fake person, based at least on the category received from the first classification module 215 and / or on the category received from the face detection module 211.
[0118] According to some other implementations, the decision module 216 is configured to determine a final classification or final decision based on the category received from the first classification module 215 and / or based on the category received from the face detection module 211 and / or based on information received from the facial key point detection module 212 and / or based on the category received from the second classification module 218, as will be described below.
[0119] In practice, according to these implementations, after face detection and keypoint detection, the keypoint detection module 212 sends the thermal image 201, along with optionally the coordinates of the keypoints, to a large face cropping module 217, which is configured to crop the thermal image 201 to obtain a second cropped thermal image. The second cropped thermal image is cropped over a larger area than the first cropped thermal image, which is a canonical face alignment. For example:
[0120] - Standardized face alignment includes only the face of the target person in the first cropped thermal image and has no background, or the background is less than 5%;
[0121] - Large face cropping includes more than 5% of the target person's face and background in the second cropping heat image, for example, more than 30% of the background.
[0122] The large face cropping module 217 is configured to send the second cropped thermal image to the second classification module 218.
[0123] Alternatively, the large face cropping function is performed by the key point detection module 212, and in this embodiment, the attack detection device 210 does not include the large face cropping module 217. Therefore, the second cropped thermal image is sent from the key point detection module 212 to the second classification module 218.
[0124] The second classification module 218 is configured to receive thermal images (such as a second cropped thermal image) as input and is configured to determine one of two categories, “live” or “deceive”.
[0125] To this end, the second classification module 218 can implement a fifth model, which is configured to receive thermal images (such as a second cropped thermal image) as input and is configured to classify target persons in the input thermal images into the category of "live" or "deceived".
[0126] There are no restrictions on the fifth model; it can be analytical or it can be obtained based on machine learning.
[0127] For example, the fifth model obtained based on machine learning could be a neural network, such as a deep neural network, configured to extract features. The output layer of the fifth layer is configured to predict the category of liveness or deception based on features extracted from the input thermal image.
[0128] The fifth model can be trained through supervised learning based on a fifth set of training data, which includes reference thermal images labeled with the categories of liveness or deception.
[0129] The second classification module 218 is configured to send the determined category to the decision module 216.
[0130] It should be noted that, according to the present invention, modules 217 and 218 are optional.
[0131] The decision module 216 is configured to determine the final decision based on the category received from the face detection module 211 and / or further based on the category received from the face detection module and / or the category received from the second classification module 218.
[0132] To this end, decision module 216 may apply a pre-determined fusion function to determine a PAD score based on at least one received category. A final decision may then be determined based on the PAD score. For example, the final decision may include binary information indicating whether the target person included in thermal image 201 is “alive” or “deceived.”
[0133] The final decision may also include a PAI identifier, which identifies the tool used by the target person during PA in the case of binary indication "deception". When a category is received from the first classification module 216, and at least another category is received from the face detection module 211 and / or from the second classification module 218, the PAI identifier may be determined based on the PAD score or based on a predetermined rule applied to the category.
[0134] The final decision determined by the decision module 216 may be sent to an external entity via the output interface 231 and / or may be stored in the memory 220 of the attack detection device 210.
[0135] For example, the final decision may be sent to a facial recognition device along with the thermal image 201. If the thermal image 201 is determined to be a live person, the facial recognition device may compare the thermal image 201 or its features with features from a reference image in a database to identify the target person in the thermal image 201.
[0136] Alternatively, facial recognition can be performed by presentation attack detection device 210. For example, presentation attack detection device 210 may store a database of reference images associated with corresponding identity information of a person. If the final decision is deception, thermal image 201 may be discarded by presentation attack detection device 210. If the final decision is liveness, presentation attack detection device 210 may compare thermal image 201 or the features of thermal image 201 with the reference images to determine the identity information of the target person in thermal image 201.
[0137] Facial recognition operations are well-known and will not be described further in this specification.
[0138] Therefore, the present invention allows for the detection of a wide variety of PAs, including full-face latex or silicone masks, by using a facial detector applied to a facial thermogram and then using a domain converter.
[0139] In practice, thermal imaging cameras struggle to capture faces when using PAIs such as printed facial images, 2D or 3D masks, smartphones, or tablets, because radiative information is obstructed in thermal images that display thermally uniform shape and temperature, as will be described below. Figure 4 and Figure 5 As described.
[0140] Therefore, in this case, the face detection module 211 fails to detect a face and acts as a PA detector: the category "no face" sent to the decision module 216 can be interpreted as PA based on the PAI listed above.
[0141] For more advanced PAIs, such as those based on full-face latex or silicone masks, the face is detected by the face detection module 211, and the thermal image is further processed by other modules: then, the domain conversion module 214 generates unrealistic images that include perceptual anomalies that can be easily detected by the first classification module 215.
[0142] As referenced above Figure 1 The explanation is that when a full-face mask is worn, the mask warms up upon contact with the skin compared to when the target person first wears a cooler mask, thus distributing heat energy more evenly. Although heat is distributed across the face, higher energy escapes from specific areas such as the eyes, mouth, and neck. This variation in facial heat distribution becomes a compelling clue after domain conversion by domain conversion module 214, since the eye and mouth areas are relevant biometric identification features. Therefore, changes in such features interfere with domain conversion module 214, resulting in an unrealistic face classified as deceptive by first classification module 215.
[0143] In addition, as explained above, the domain conversion module 214 can also receive thermal images with different resolutions, which allows PAD to be performed passively without requiring cooperation from the target personnel.
[0144] The first to fifth models described above can be stored in the modules that implement them respectively, or alternatively in the memory 220 of the attack detection device 210.
[0145] Figure 3 Examples of generative adversarial networks (GANs) 300 trained to obtain a third model used by the domain transformation module 214 according to some embodiments of the present invention are shown.
[0146] The third model, also known as "ANYRES," enables simultaneous facial super-resolution and thermal-to-secondary spectral domain (such as the visible domain) transformations while preserving identity, and is robust to any low-resolution thermal input. The benefits of this simultaneous process help avoid both accumulated errors and artifact generation, and simultaneously bridge the modal and resolution gaps. See [reference to specific details] Figure 3 Following the training process described below, the blurry, thermal, and low-resolution (LR) facial image 306 is transformed into a clear, realistic, high-resolution visible facial image 320.
[0147] The designed third model has the advantage of preserving consistent biometric features across low / high resolution spatial and thermal / visible spectral dimensions. Furthermore, the implementation is adapted to real-world scenarios because the distance between the human and the camera is random during operational applications, thus allowing for the creation of multi-scale LR thermal images (depending on the acquisition distance). Unlike known techniques, where the resolution is fixed as the input, ANYRES emphasizes its ability to operate at any input resolution, from low to high. Figure 3 In the process, the thermal image ranges from higher resolution to lower resolution across layers 308a to 308e of the encoder 302 of the generative adversarial network 300. The first layer 308a is arranged to receive the thermal image 306 as input.
[0148] Generally, in some implementations, 308a to 308e are coding layers rather than images. Each of 308a-308e is no longer a corresponding image. 308a will be a collection of graphs (which are also referred to as images, but at this stage, the preferred technical term is "graph"). For example, 308a may have 16 graphs, 308b may have 32 graphs, and so on. The increased number of graphs is not mandatory, but it is usually the whole.
[0149] GAN 300 also includes a decoder 304, which is arranged to synthesize a visible image based on the graph received from the last layer 308e of decoder 302.
[0150] Following the training phase, encoder 302 and decoder 304 together form a third model (labeled 301), which can be implemented by domain transformation module 214 during the current phase. The third model 301 is trained as the generator part of GAN 300.
[0151] As shown by using decoder 304, the decoded synthetic visible output "images" or layers 310a-310d (again, as explained with respect to 308a-308e, will be a collection of images rather than a single image) range from higher resolution to lower resolution from 310a to 310d. In some embodiments, further detailed below, the third model performs the step of skip connection 312 between the encoded infrared facial image and the decoded high-resolution visible facial image. In some embodiments, the third model also performs the steps of squeezing and excitation 314 between the encoded infrared facial image and the decoded high-resolution visible facial image. In some embodiments, GAN 300 also uses a reference visible image 316 to consider the loss function used to train the third model during the training phase.
[0152] In GANs, the discriminator part is responsible for receiving a pair of images and determining which of these images is real and which is synthetic.
[0153] Therefore, the GAN also includes a discriminator 324, which implements one or more discriminator functions and is designed to determine which of the synthesized visible image 320 and the corresponding reference image 316 is real. If the discriminator 324 successfully guesses that the visible image 320 is synthesized, at least one change is made in a layer or connection to improve the process of synthesizing the visible image 320, and the training process is repeated by providing the discriminator 324 with a new pair of images.
[0154] Furthermore, an additional loss function 322, which will be described below, may be used during the training process, and the third model 301 is configured to minimize the additional loss function 322 or a combination of these functions, and is configured to prevent the discriminator 324 from distinguishing real visible images from synthetic images.
[0155] Figure 4 An example of a discriminator 324, which can be used during the training phase of a domain-transformed third model according to some embodiments of the present invention, is shown.
[0156] The discriminator 324 may include a global discriminator function 404 and a local discriminator function 406. While the global discriminator 404 helps generate an overall correct identity, the local discriminators 406, named L1, L2, L3, and L4, located on the eyes, nose, and mouth (in the reference visible image 316 and the composite visible image 320, respectively), are designed to focus on the details of generating cross-spectral biometric features and to help improve the identity by focusing on more biometrically relevant parts of the face.
[0157] Therefore, the discriminator 324 attempts to determine which of the input images 316 and 320 is real and which is synthetic.
[0158] For illustrative purposes, details of the problem formulas, layers, loss functions, and discriminators used during the training phase of the third layer, which serves as the generator of GAN 300, are given below.
[0159] We consider a high-resolution space with dimension m×n, combined with a visible facial image x. vis ∈R m×n The visible domain V and the thermal facial image x thm ∈R m×n The thermal domain T.
[0160] In the following text, x vis A reference visible image 316 is specified to be paired with the synthesized visible image 320 in the third set of training data, and the synthesized visible image is labeled as x. SR vis .
[0161] The third model 301 includes both a domain transformation stage and a super-resolution stage.
[0162] In the domain transformation stage, image-to-image transformation is performed by learning an end-to-end non-linear mapping between the thermal spectrum and the visible spectrum (denoted as Θ t→v ). This is formalized as follows
[0163]
[0164] Therefore, Θ t→v is a function that synthesizes the thermal face image 306 into a realistic synthetic visible face image 320x in the high-resolution space SR vis of.
[0165] In the super-resolution stage, given the embedding of Equation (1) above, the third model 301 or network 301 encapsulates super-resolution scalability as a simultaneous task. Thus, during training, the third model 301 learns a conditional generation function, where the thermal low-resolution face image is also enhanced to the high-resolution scale, thereby giving the synthetic visible image 320.
[0166]
[0167] The goal is a domain transformation that is robust to any low-resolution thermal input and aims to learn a unified function that produces a higher-resolution super-resolved visible image with rich semantic and identity information when applied to any low-resolution thermal image x LR thm . In this context, the third model according to the present invention simultaneously learns the global interaction between domain transformation and resolution scalability by enriching Equation (1) with Equation (2).
[0168] GAN 300 is trained on a third set of training data, which includes thermal images associated with reference images in the second spectral domain for all scale factors 0 < r ≤ m.
[0169] The third model 301 is based on a U-shaped pyramid architecture network and thus naturally relies on multi-scale analysis. The overall architecture is in Figure 3The following is an example. In one specific implementation, the U-Net architecture is used for its efficiency, where the generator 301 or third model 301 consists of encoder-decoder structures 302-304, which have skip connections 312 between the domain-specific encoder 302 and decoder 304. Considering the significant differences between images generated from LR and high-resolution (HR) spaces, GAN 300 further utilizes a squeeze and excitation (SE) block 314, which acts as a gate modulator after each skip connection. This strategy allows for flexible control through channel-wise relationships and balances the encoded features with the decoded super-resolution features.
[0170] During training, batches of 301 low-resolution thermal images 306 with a wide range of scale factors r are generated simultaneously. It should be noted that, in the extreme case considering only one low-resolution scale, with a fixed r, the model will be able to... Super-resolution of thermal images is performed in a space of scale m×n (i.e., a fixed low-resolution input that is different from any low-resolution input). A model trained using a single scaling factor is called a single-resolution model, while a model trained using several scaling factors is called a multi-resolution model.
[0171] The encoder 302 can extract multi-resolution features in parallel and repeatedly fuse these multi-resolution features during learning to generate a high-quality SR representation 320 with rich semantic / identity information.
[0172] Given an LR hot input image 306, layer H0 transforms the LR hot input image into a high-dimensional feature space:
[0173]
[0174] Here, H0 refers to the composite function of two consecutive convolution-batch normalization-ReLU layers. Then, generator 301 applies a series of operations:
[0175]
[0176] Where F i This represents the intermediate feature map of the encoding after the i-th operation, where for all i∈[1,K], Here, H i It is the same composite function defined in equation (3), and pool represents the max pooling operation, in which the most prominent features of the previous feature map are preserved. This is implemented for layers 308a to 308e.
[0177] Decoder 304 aims to transform the high-dimensional feature space into a super-resolution image 320 in the visible domain. Therefore, the task of generating the super-resolution image 320 begins from the deep layer (U bottleneck) with the following equation (5):
[0178]
[0179] Then, in ascending order, for all Layers 310a to 310c implement:
[0180]
[0181] The decoding process ends with the generation of a visible image 320 through a convolutional hyperbolic tangent layer, as given by the following equation (7):
[0182]
[0183] Although S refers to the factorial upsampling operation, followed by a convolution-batch normalization-ReLU layer, C will come from all channels of the skip connection 312. (See 311a) and upsampling S i Layers (see 309a) are cascaded. Finally, G i This represents the intermediate feature map of the decoding after the i-th operation prior to squeezing and stimulating SE 314 (e.g., see 310a).
[0184] The discriminator function implemented by discriminator 324 of GAN 300 is described in detail below.
[0185] As explained above, the discriminator includes components named Dis. 全局 and Dis 局部 The generator employs a global discriminator and a local discriminator. The former helps the generator synthesize photorealistic HR images in the visible spectrum, while the latter (in some implementations) focuses on the fine details of each individual face and benefits from the inherent local focus on capturing faithful biometric features during generation.
[0186] Global discriminator: In some implementations, a multi-scale discriminator is used, which enables the generation of realistic images with refined details. Figure 4 The global discriminator 404 in the image is responsible for passing the super-resolution or synthetic visible image 320 x SR vis Compared with the reference visible image x vis 316 are separated and binary classification is performed.
[0187] Local discriminator: To synthesize realistic semantic content for biometrics, implementations can focus on discriminative regions related to identity information, such as the eyes, nose, and mouth. These regions of interest are defined by the image x. vis With x SR vis The same cropping area between ( Figure 4 (As shown) indicates that they are respectively labeled as , where i∈[0,4]. Each independent discriminator focuses on the fine details of each individual face and benefits from the inherent local focus on capturing faithful biometric features during generation.
[0188] The adversarial learning process of the third model 301 can be further enhanced through efficient combination of objective functions, and paves the way for controlling the synthesis process at both the pixel and feature levels.
[0189] On one hand, adversarial losses, including global and local ones implemented by the discriminator 324, are responsible for making the generated samples realistic and indistinguishable from real images in the target domain. On the other hand, additional loss functions 322 (such as L1 loss), described below, drive the spectral transformation, while perceptual loss, identity loss, and attribute loss are used at high-level features to influence the perceptual rendering of the image, preserve biometric identity-specific features, and implement consistent age-gender reconstruction, respectively. All the combined loss functions contribute realism during the spectral transformation and avoid blurring introduced by any low-scale resolution from the thermal image input.
[0190] Regarding the adversarial loss, the image generated by the third model through equation (1) must be realistic. Therefore, the goal of the third model is to maximize the probability that the discriminator 324 makes an incorrect decision. On the other hand, the goal of the discriminator 324 is to maximize the probability of making a correct decision, i.e., to effectively distinguish between real and synthetic images. Global loss function and local loss function It is part of adversarial training and is defined as follows:
[0191]
[0192] The additional loss function 322 is described below.
[0193] Setting conditions for the spectral distribution is advantageous for generating images within the target spectrum. Conditional loss L cond (Or L1 loss) is defined in equation (9) as follows:
[0194]
[0195] Perceived loss L PThis method influences perceptual rendering of images by measuring high-level semantic differences between synthesized and target facial images. It reduces artifacts and enables the reproduction of realistic details. P Defined as follows
[0196]
[0197] in This represents features extracted from VGG-19 pre-trained on ImageNet.
[0198] Identity loss L I The identity of the facial input is preserved, and a pre-trained ArcFace recognition network is used to extract facial feature embeddings. Then, a cosine similarity metric is used to provide the identity loss function.
[0199]
[0200] Attribute loss L A To prevent attribute shifts during domain transformation. Specifically, age and gender information are subtly unavailable on thermal images. While age provides apparent information, gender is identity-dependent. Therefore, apparent age loses L... 年龄 A and gender loss L 性别 A Defined as follows:
[0201]
[0202] in 年龄 and 性别 It is a pre-trained model based on the DeepFace facial attribute framework analysis. Then, the attribute loss is expressed as the following equation (14):
[0203]
[0204] Finally, the overall loss function used for the proposed third model 301 can be a combination of the aforementioned loss functions or any combination thereof.
[0205] Figure 5 Thermal images 501 to 503 of a presentation attack based on a 2D PA tool are shown.
[0206] Thermal images 501 and 502 show a presentation attack in which a target person holds a piece of paper representing the face of another person.
[0207] Thermal image 503 shows a presentation attack where the target person is holding a tablet computer on which the face of another person is displayed.
[0208] As can be seen, acquiring the thermal image makes it easy to detect the two-dimensional PA, because the face detection unit 211, which receives thermal images 501 to 503, does not detect a face. Then, the decision module 216 can detect the PA based on the two-dimensional PA image.
[0209] Figure 6 A PA based on a paper 3D face mask is shown. In thermal image 601, a person begins to wear the paper 3D face mask, which is therefore cold. In thermal image 602, the same person is shown wearing the same paper 3D face mask after a certain amount of time.
[0210] Acquiring thermal images enables the display of high contrast between areas of hair, eyes, and face covered by the mask. This high contrast can be detected by the second classification module 218 that processes the thermal images, but can also be detected by the first classification module 215 after domain transformation.
[0211] Therefore, the decision module 216 can detect PA based on a 3D paper mask.
[0212] For more advanced attacks, such as those based on Figure 1 The attack shown is a full-face latex or silicone mask. Using thermal imaging combined with domain transformation, the first classification module 215 can detect that the target person is deceiving.
[0213] Therefore, the decision module 216 can detect PA based on a full-face latex or silicone mask.
[0214] Figure 7 The steps of a method for determining viability according to some embodiments of the present invention are shown.
[0215] The method includes a preliminary stage 700 for constructing one or more models stored in the previously described presentation attack detection device 210.
[0216] At step 710, the first set of training data is obtained, as previously explained.
[0217] The first model can be trained in step 711 based on the first set of training data. Alternatively, the first model may be obtained not through machine learning, but through analysis.
[0218] Then, at step 712, the obtained first model is stored in the face detection module 211 or in the memory 220 of the presentation attack detection device 210.
[0219] As explained above, the first model is configured to receive thermal images and to determine whether a face is detected in the thermal images.
[0220] At step 720, a second set of training data is obtained, as previously explained.
[0221] The second model can be trained at step 721 based on a second set of training data. Alternatively, the second model may be obtained analytically rather than through machine learning.
[0222] Then, at step 722, the obtained second model is stored in the key point detection module 212 or in the memory 220 of the presentation attack detection device 210.
[0223] As explained above, the second model is configured to receive a thermal image in which a face has been detected, and is configured to determine at least one key point in the thermal image.
[0224] At step 730, a third set of training data is obtained, as previously explained.
[0225] The third model can be trained in step 731 based on the third set of training data. (Referring to the above...) Figure 3 and Figure 4 At that time, a detailed example of training a third model had already been described.
[0226] Then, at step 732, the obtained third model is stored in the domain conversion module 214 or in the memory 220 of the presentation attack detection device 210.
[0227] As explained above, the third model is configured to receive thermal images and output corresponding images in the second spectral domain, such as visible images.
[0228] At step 740, a fourth set of training data is obtained, as previously explained.
[0229] A fourth model can be trained at step 741 based on a fourth set of training data. Preferably, the fourth model can be a neural network, such as a deep neural network. Alternatively, the fourth model is obtained not through machine learning, but through analysis.
[0230] Then, at step 742, the obtained fourth model is stored in the first classification module 215 or in the memory 220 of the attack detection device 210.
[0231] As explained above, the fourth model is configured to receive images in the second spectral domain, such as visible images, and is configured to output the category of liveness or spoofing based on the visible images.
[0232] At step 750, the fifth set of training data is obtained, as previously explained.
[0233] A fifth model can be trained at step 751 based on a fifth set of training data. Preferably, the fifth model can be a neural network, such as a deep neural network. Alternatively, the fifth model is obtained not through machine learning, but through analysis.
[0234] Then, at step 752, the obtained fifth model is stored in the second classification module 218 or in the memory 220 of the attack detection device 210.
[0235] As explained above, the fifth model is configured to receive thermal images and output the category of liveness or deception based on the thermal images.
[0236] The initial phase can be implemented in a device that includes processing power and memory, which is external to the attack detection device 210.
[0237] The method also includes a current phase 800 implemented by the attack detection device 210 according to some embodiments of the invention.
[0238] At step 801, the attack detection device 210 receives at least one thermal image 201 from the thermal imaging camera 200 via the input interface 230.
[0239] At step 802, the face detection module 211 applies the first model to the received thermal image to determine whether a face is detected in the thermal image 201.
[0240] At step 803, if a face is detected, the face detection module 211 forwards the thermal image 201 to the keypoint detection module 212. As explained above, the face detection module 201 and the keypoint detection module 212 can be the same module, so that the thermal image 201 is not forwarded between two different modules. If no face is detected, the method proceeds to step 808.
[0241] Furthermore, at step 803, the face detection module 211 sends the categories in "face" and "no face" to the decision module 216.
[0242] At step 804, the keypoint detection module 212 determines whether at least one keypoint is detected in the thermal image 201 by applying the second model to the thermal image 201, as previously described. Information indicating whether at least one keypoint is detected is sent to the final decision module 216.
[0243] At step 805, the face alignment module 213 performs canonical face alignment on the thermal image 201 based on at least one key point detected by the key point detection module 212 to obtain a first cropped thermal image.
[0244] At step 806, the domain conversion module 214 converts the first cropped thermal image in a second spectral domain (such as in the visible domain) by applying a third model, as previously explained.
[0245] At step 807, the first classification module 215 applies the fifth model to the image in the second spectral domain to determine the category of “deception” and “liveness”, and sends the determined category to the decision module 216.
[0246] At step 808, decision module 216 determines the final classification or final decision of the target person in thermal image 201 as either a live person or a fraud, based at least on the category received from first classification module 215 and optionally on the category received from face detection module 211. According to some embodiments, the final decision may further be based on the category received from second classification module 218, as described below at steps 810 and 811.
[0247] At step 809, the liveness determination device 211 may send the final decision to an external entity via output interface 231.
[0248] At step 810, the large face cropping module 217 crops the thermal image 201 to obtain a second cropped thermal image.
[0249] At step 811, the second classification module 218 applies the fifth model to the second cropped thermal image to determine the category of deception or liveness, as previously described. The second classification module 218 also sends the determined category to the decision module for making a final decision at step 808.
[0250] Therefore, this invention allows PA detection to be performed using only a thermal imaging camera 200. Unlike the visible spectrum, thermal radiation highlights PAI. For example, even when the tool is heating up, a deception mask is characterized spectrally via thermal interference. Thus, this invention utilizes thermal interference, measurement, and visualization to determine liveness with the aid of continuous face detection and domain transformation.
[0251] Face detection serves as an initial PAD filter to detect real faces under broad conditions, while domain transformation is dedicated to converting the hot face image into a synthetic face in another domain while preserving identity. The resulting synthetic face image in the second spectral domain then feeds relevant information into a fourth model used by the first classification module for PAD.
[0252] Due to the unnatural distribution of facial temperature, the deceptive facial image will highlight the inconsistent perceptual transformation image (in the second spectral domain) with numerous artifacts. Furthermore, the invention is adaptable to real-world scenarios because face detection can operate under a wide range of conditions, and according to some embodiments of the invention, the domain transformation accepts thermal images with various resolutions / qualityes.
[0253] This makes the present invention a robust PAD solution for all types of deception, including heated masks, while providing a solution for determining liveness without additional sensors.
[0254] This invention is based on real-world PA (Power Amplifier) and benefits from multi-channel decision-making, which avoids errors and thus achieves robust PAD (Power Amplifier Device) capabilities. Furthermore, thermal imaging cameras are being deployed more and more frequently.
[0255] The example implementations have been described in sufficient detail to enable those skilled in the art to emulate and implement the systems and processes described herein. It is important to understand that implementations may be provided in many alternative forms and should not be construed as limited to the examples set forth herein.
Claims
1. A method for detecting presentation attacks, the method comprising: - Receive (801) at least one thermal image (201) from thermal imaging camera (200); - Determine (802) whether the face of the target person is detected in the thermal image; -If the face of the target person is detected in the thermal image, perform (806) domain transformation based on the thermal image to obtain an image in a second spectral domain, which is different from the thermal domain; -Based on the image in the second spectral domain, determine (807) a first category indicating whether the target person is alive or deceiving; -A final decision (808) representing the liveness of the target person is determined based on the first category determined and / or based on whether the face of the target person is detected in the thermal image.
2. The method of claim 1, further comprising performing (805) canonical face alignment of the thermal image (201) to obtain a first cropped thermal image, wherein the domain transformation is performed (806) on the first cropped thermal image to obtain the image in the second spectral domain.
3. The method according to claim 1 or 2, further comprising detecting (804) at least one key point of the face of the target person if (802) the face of the target person is detected in the thermal image (201), wherein the domain transformation is performed if the face of the target person is detected and if at least one key point of the face of the target person is detected, and wherein the final decision is further based on whether at least one key point of the face of the target person is detected in the thermal image.
4. The method according to claims 2 and 3, wherein the standardized facial alignment is based on at least one key point detected on the face of the target person.
5. The method according to any one of the preceding claims, wherein the domain transformation is performed (806) by applying a model (301) configured to determine an image in the second spectral domain for a corresponding thermal image received as input.
6. The method of claim 5, wherein the model (301) is arranged to perform super-resolution to receive a thermal image with a first resolution as input, and is arranged to output an image with a second resolution in the second spectral domain that is higher than the first resolution.
7. The method according to claim 5 or 6, wherein the model (301) has an encoder-decoder structure based on a pyramidal architecture.
8. The method according to any one of claims 5 to 7, the method comprising a preliminary step (731) of training the model (301) using a generative adversarial network (300), wherein the generative model is trained by a discriminator (324) using at least one adversarial loss function (404; 406).
9. The method of claim 8, wherein the model (301) is trained by the discriminator (324) using a global adversarial loss function (404) and at least one local adversarial loss function (406).
10. The method according to claims 8 and 9, wherein the model (301) is further trained based on at least one additional loss function (322) among the L1 loss function, the perceptual loss function, and the identity loss function.
11. The method according to any one of claims 8 to 10, wherein the model (301) is trained by supervised learning on a training dataset comprising a plurality of thermal images with a first resolution associated with a reference image having a second resolution in the second spectral domain, the second resolution being higher than the first resolution.
12. The method according to any one of the preceding claims, wherein the second spectral domain is the visible domain.
13. The method according to any one of the preceding claims, the method further comprising, in the case that the face of the target person is detected in the thermal image, determining (811) a second category indicating whether the target person is alive or deceiving based on the thermal image (201); and wherein the final decision is further based on the second category.
14. The method according to claim 13, wherein the first category is determined (807) by a first classification module (215) based on a corresponding classification model, and the second category is determined (811) by a second classification module (218).
15. A computer program comprising instructions arranged to implement the method according to any one of the preceding claims when executed by a processor.
16. A presentation attack detection device (210), the presentation attack detection device comprising: - Receive interface (230), the receive interface being used to receive at least one thermal image (201); - Face detection module (211), the face detection module being configured to determine whether the face of a target person is detected in the thermal image; - Domain conversion module (214), the domain conversion module is configured to perform domain conversion based on the thermal image to obtain an image in a second spectral domain, which is different from the thermal domain; - A first classification module (215), configured to determine a first category indicating whether the target person is alive or deceiving based on the image in the second spectral domain; - Decision module (216), which is configured to determine a final decision representing the liveness of the target person based on the determined first category and / or based on whether the face of the target person is detected in the thermal image.