Method, device and system for detecting facial modification

EP4740184A1Pending Publication Date: 2026-05-13BUNDESDRUCKEREI GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
BUNDESDRUCKEREI GMBH
Filing Date
2024-06-20
Publication Date
2026-05-13

AI Technical Summary

Technical Problem

The risk of facial manipulation during self-enrollment in facial recognition systems, where individuals may deceive the system by presenting altered or falsified images, is a challenge in ensuring the authenticity of identity verification.

Method used

A method that combines color and thermal images to detect facial manipulation by aligning and combining them, using a convolutional neural network to determine probability values for specific types of manipulations, and classifying based on these values to determine if facial manipulation has occurred.

Benefits of technology

This approach significantly enhances the detection of facial manipulation by utilizing the unique heat distribution patterns in thermal images to identify and classify attempts to alter or obscure facial features, thereby ensuring the authenticity of the identity verification process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure DE2024100549_09012025_PF_FP_ABST
    Figure DE2024100549_09012025_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to a method for detecting facial modification, the method comprising the following steps carried out by means of a computer (1): receiving at least one pair of images with a colour image of a face and a heat image, captured simultaneously, of the face; adapting the colour image and the heat image so that the face in the colour image is congruent with the face in the heat image; generating a combined image by combining the adapted colour image and the adapted heat image; determining from the combined image, for one or more classes indicating a type of modification, a probability value in each case that specifies a probability that a modification is present for the corresponding class; and classifying, on the basis of the probability values, whether facial modification is present.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Method, device and system for detecting facial manipulation

[0002] The invention relates to a method for detecting facial manipulation, a device for detecting facial manipulation, a system for detecting facial manipulation, and a computer program product comprising instructions that cause a correspondingly designed device to carry out the method for detecting facial manipulation.

[0003] A person's face is crucial to their identity. A visual representation of the face of a person to be identified is often captured using a color image for passport or ID photos. A color image can, for example, be an RGB color image, a CMYK color image, or a monochrome image. These images should be authentic, i.e., reflect the person's real face, allowing reliable conclusions to be drawn about the person's identity, for example, in personalized ID cards.

[0004] Facial recognition methods are also known in which a camera captures an image of a person's face in the visible wavelength range. The resulting image provides a representation of the person's face. A person can be uniquely identified by analyzing a geometric arrangement of characteristic facial features such as the eyes, nose, or mouth, as well as their position, relative distance, and location within the facial area—that is, an area on the person's frontal head that is specific to each individual.

[0005] A corresponding image of the person's face can, for example, be taken on site by an employee of a security service provider, in which the employee takes an image of the face of the person to be identified on site.

[0006] Recently, there has been a trend toward capturing a picture of a person's face without the involvement of another person, such as a service provider employee. For this purpose, a so-called self-enrollment system can be implemented, in which a person to be identified takes a picture of their face independently, without the presence of another person. A possible identification process or facial recognition process can then be carried out automatically by software based on the captured image of the person's face.

[0007] Since the image of the face of the person to be identified is now taken without the presence of an authorized person, there is a possibility that the person to be identified could manipulate the system.

[0008] In the worst case, the person to be identified tries to deceive the system by presenting the system with a different identity than their own by presenting the system with a different or at least partially falsified face.

[0009] Such attempts at deception may, for example, consist of the person to be identified bringing an image of another person at least partially into the field of view of the camera, putting on a mask, holding a printed or digital image with the image of another or even non-real person into the field of view of the camera or covering certain facial features, for example with sunglasses, an eye patch or similar.

[0010] Through such deception attempts, the image captured by the camera shows a different or at least partially falsified face of the person to be identified.

[0011] There is therefore a risk that the image taken of the face of the person to be identified will show a different or falsified face of the person and thus create a false identity of the person to be identified.

[0012] Document US 2021 / 110018 A1 describes systems and methods for using machine learning for image-based forgery detection, including a method for retrieving an input dataset containing multiple images of a biometrically authenticated subject captured by multiple cameras of a camera system. The method further includes inputting the input dataset to a trained machine learning module and processing the input dataset using the machine learning module to obtain a forgery detection result for the biometric authentication subject from the machine learning module. The method also includes outputting the forgery detection result for the biometric authentication subject.

[0013] In the document Chen, X. et al., IR and visible light face recognition, Computer Vision and Image Understanding, 2005, Vol. 99, No. 3, pp. 332-358, the results of several large-scale studies on face recognition using visible light and infrared imaging in the context of principal component analysis are presented.

[0014] Document DE 10 2019 135 107 A1 relates to a device and a method for determining a biometric feature of a person's face. Document US 7,421,097 B2 describes facial identification using three-dimensional modeling.

[0015] Summary

[0016] Against this background, one object of the invention is to detect facial manipulation when taking a picture.

[0017] The problem is solved by a method for detecting facial manipulation, which comprises the following steps carried out by a computer: receiving at least one image pair comprising a color image of a face and a simultaneously recorded thermal image of the face; adjusting the color image and the thermal image such that the face in the color image is congruent with the face in the thermal image; generating a combined image by combining the adjusted color image and the adjusted thermal image; determining, on the combined image, for one or more classes indicating a type of manipulation, a probability value which indicates a probability that manipulation is present for the corresponding class; and classifying, based on the probability values, whether facial manipulation is present.The invention is based on the inventors' finding that a combination of a color image in the visible wavelength range and a thermal image in the infrared wavelength range, which indicates heat distribution, can significantly increase security. A thermal image can be, for example, a FLIR image.

[0018] Both the color image and the thermal image taken at the same time show a similar section, preferably the same section, with the face of a person to be identified.

[0019] The color image and the thermal image are adjusted so that the face in the color image and the face in the thermal image are congruent. Congruent means that two objects are congruent, i.e., that the face is congruent in both images. In other words, all points of the face in the color image correspond to the points of the face in the thermal image. The face in the color image and the face in the thermal image can be approximately congruent. Congruent also includes smaller deviations of the faces in the range of 0.00001%-10%, for example 0.01, 0.1, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10% or in other words in the range of 0.0000001 - 10 mm, for example 0.01, 0.1, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 mm.

[0020] The adjusted thermal image and the adjusted color image are combined, for example, by stacking (stack operation). The combined image contains the face, with the information from the color image linked to the information from the thermal image for each available point of the face in the combined image. Thus, location-dependent information in the visible wavelength range is linked to information in the infrared wavelength range. The combined image can, for example, take the form of a tensor in which the information from the color image and the thermal image is combined. When the color image and the thermal image are combined by stacking, the combined image can also be referred to as a stacked image.

[0021] The combined image can help to better identify whether facial manipulation has occurred, as a face usually has a specific heat distribution.

[0022] When a person uses a representation of another person's face, such as a print or a digital representation through a screen, or obscures certain areas of their own face, the information of the thermal image in the combined image changes because the heat distribution of the captured face is changed compared to the unaltered face by the facial manipulation.

[0023] These different types of manipulation have a specific heat distribution. Thus, a type of manipulation can be assigned a class with a specific heat distribution. For example, a first class can include holding a printed image with a reproduction of the face up to the camera. The printed image has a first location-dependent specific heat distribution. A second class can include putting on a mask. The reproduction of a face with a mask on has a second location-dependent specific heat distribution. A third class can include, for example, putting on glasses with image elements. The reproduction of a head with such glasses on has a third location-dependent specific heat distribution.Such location-dependent specific heat distributions can be determined for different types of manipulations, whereby the number of types of manipulations is not limited to the examples mentioned and a multitude of such classes, ie types of manipulations, can be formed.

[0024] For each type of manipulation, a probability value can be determined from the combined image, indicating the probability that a type of manipulation has occurred.

[0025] By determining the probability value for each class, i.e. a type of manipulation, it can be classified whether manipulation has occurred.

[0026] This solves the task mentioned at the beginning by detecting or classifying whether facial manipulation has occurred.

[0027] Further advantageous aspects and embodiments of the invention are described below.

[0028] According to one aspect, adjusting the color image and the thermal image comprises the steps of: determining characteristic facial features of the face in the color image; scaling dimensions of the color image such that a distance between two of the characteristic facial features of the face in the color image has a predetermined value; perspective transforming the thermal image using a coordinate transformation such that a perspective of the thermal image and a perspective of the color image match; and scaling dimensions of the thermal image to the dimensions of the scaled color image such that the face in the scaled thermal image is congruent with the face in the scaled color image.

[0029] A coordinate transformation may preferably be an affine mapping.

[0030] This aspect allows the color image and the thermal image to be adjusted particularly precisely such that the face in the color image is congruent or substantially congruent with the face in the thermal image.

[0031] Typically, characteristic facial features are more pronounced and identifiable in the visible range. Furthermore, a color image usually has a higher resolution than a thermal image. This allows the characteristic facial features to be identified more accurately in the color image than in the thermal image. The thermal image is thus adjusted to the color image based on the information in the color image so that the face in the thermal image is congruent with the face in the color image. The distance between two of the characteristic facial features of the face can be a Euclidean distance between the two eyes of the face and can be, for example, on the order of 60 or 100 pixels.

[0032] Since the color image is typically captured by a first camera and the thermal image by a second camera, which is spaced apart from the first camera and thus has a slightly different perspective than the first camera, the thermal image is transformed using a perspective coordinate transformation, in particular an affine mapping, so that the perspective of the first camera, or the color image, and the second camera, or the thermal image, match. This improves the accuracy of the method and compensates for mechanical inaccuracies between devices.

[0033] Based on the scaling of the color image, the thermal image can be scaled to the same dimensions as the scaled color image. A dimension can include, for example, the horizontal and vertical dimensions of the corresponding image, but also optionally the pixel resolution of the corresponding image.

[0034] After adjusting the color image and the thermal image, the face in the color image and the face in the thermal image are essentially congruent and both images have the same cropping.

[0035] According to one aspect, the method comprises: cropping the scaled color image based on the characteristic facial features such that the face is arranged at a predetermined position and with a predetermined size in the color image; and cropping the scaled thermal image analogously to the cropping of the color image such that the cropped thermal image and the cropped color image have the same section.

[0036] Based on the characteristic facial features determined in the color image, the scaled color image and the scaled thermal image can be cropped so that the face in these images is positioned at a predetermined position within the corresponding image. The cropping can, for example, be designed so that the face in the corresponding color image and thermal image complies with ICAO regulations.

[0037] This ensures that other areas of the body, such as the shoulder or torso and the background, are not evaluated to detect facial manipulation, which could distort the evaluation. Furthermore, it ensures a certain degree of uniformity, which facilitates the evaluation, especially the determination of probability values.

[0038] According to one aspect, the characteristic facial features in the color image are determined using artificial intelligence.

[0039] The application of artificial intelligence has proven particularly effective in recognizing characteristic facial features. In particular, it has been shown that artificial intelligence can be trained particularly effectively to detect differences between images. One example of an artificial intelligence capable of recognizing specific facial features is the BlazeFace algorithm.

[0040] According to one aspect, the method further comprises: normalizing, by the computer, the probability value for each of the one or more classes to a normalized probability value; and wherein classifying whether facial tampering is present comprises summing the normalized probability values ​​for a plurality of classes, classifying that facial tampering is present if the sum of the normalized probability values ​​is greater than a first threshold, and classifying that facial tampering is not present if the sum of the normalized probability values ​​is less than the first threshold or less than a second threshold that is lower than the first threshold.

[0041] By normalizing the probability value for each type of manipulation or class and summing several probability values ​​for different types of manipulation, an overall probability that manipulation is present can be determined.

[0042] The overall probability can be used to classify whether facial manipulation has occurred.

[0043] The normalization of the probability values ​​to normalized probability values ​​can be carried out, for example, by a softmax function.

[0044] According to one aspect, the determination of a probability value for one or more classes indicating a type of manipulation is carried out by means of a convolutional neural network.

[0045] Such a convolutional neural network (CNN) has proven particularly effective and reliable in determining a probability value for any type of manipulation in the combined image, particularly based on the thermal information provided by the thermal image. Examples of such convolutional neural networks include the so-called Zalando CNN and the so-called MNIST CNN.

[0046] The convolutional neural network can be trained in such a way that the network determines a probability value for each type of manipulation in the form of a class.

[0047] For this purpose, a large number of images are provided for each type of manipulation, some of which represent a manipulation and some of which do not represent a manipulation attempt. For example, for a class of mask manipulation attempt, i.e. a type of manipulation in which the person to be identified puts on a mask, a certain number of images are provided in which the person is not wearing a mask and a certain number of images are provided in which the person puts on a mask. For each image, it is indicated whether it is a manipulation or not. The large number of inputs can thus be used to form certain patterns from which a probability value can be derived from a combined image, which indicates the probability of manipulation. This process is repeated for each class indicating a type of manipulation, whereby the number of classes is not limited.The convolutional neural network is therefore trained in advance for each class.

[0048] This allows probability values ​​for the different classes indicating a type of manipulation to be determined from such a combined image.

[0049] Normalizing these probability values ​​has proven particularly useful in conjunction with the probability values ​​output by the convolutional neural network, since the latter typically outputs unnormalized probability values. This results in better comparability and improved classification.

[0050] According to one aspect, the method further comprises: determining, on the combined image, an orientation of the Frankfort horizontal of a head encompassing the face relative to a horizontal and an orientation of the frontal plane of the head relative to a vertical; discarding the combined image if an angle between the Frankfort horizontal of the head and the horizontal exceeds a third threshold and / or if an angle between the frontal plane of the head and the vertical exceeds a fourth threshold.

[0051] The orientation of the head encompassing the face can be determined from the combined image. This allows selection for further processing by using only those images in which the face is captured at a specific angle. This makes the determination of the probability value for each class more precise. The orientation of the head can be determined, for example, using the method described under file number 102023 105433.3.

[0052] According to one aspect, the method comprises: capturing a color image using a first camera configured to capture light with wavelengths in the visible range; and capturing a thermal image simultaneously with capturing the color image using a second camera configured to capture wavelengths in the infrared range; generating the at least one image pair by associating the color image and the simultaneously captured thermal image.

[0053] By assigning the color image and the thermal image taken at the same time as the color image, the color image and the thermal image taken at the same time as the color image are linked to each other and thus form an image pair.

[0054] A color image is an image in the visible wavelength range. A color image can, for example, be an RGB image in the RGB color space. R stands for a red color channel, G for a green color channel, and B for a blue color channel. In some aspects, a color image can be a CMYK image in the CMYK color space. C stands for a cyan color channel, M for a magenta color channel, Y stands for a yellow color channel, and K for a black color channel. In further aspects, a color image can be a monochrome image. A color image is a pixel-based two-dimensional image, wherein the pixels are arranged in a grid. A camera designed to generate a color image records wavelengths in the visible range, i.e., in a wavelength range from 400 nm to 700 nm, using a pixel-based sensor. The first camera can thus be referred to as a color image camera. A thermal image can, for example, be a FLIR image (FLIR=Forward Looking Infrared).A FLIR image is a pixel-based, two-dimensional image, usually with a grayscale channel, where the pixels are arranged in a grid. A camera designed to capture a thermal image can also be called a thermal imaging camera. A thermal imaging camera captures wavelengths in the range of 3.5 pm - 15 pm, preferably 7 pm - 14 pm, using a pixel-based sensor. The second camera can also be called a thermal imaging camera. A thermal image can be an image in the infrared wavelength range.

[0055] The first camera and the second camera are preferably controlled such that they simultaneously capture a color image and a corresponding thermal image of the face. In other words, the thermal image and the color image can be captured at approximately the same time. "Simultaneously" within the meaning of this application means with a negligible time delay. For example, the thermal image and the color image can be captured with a time delay in the range of 1 ns - 10 ms. In some aspects, the thermal image and the color image can be captured with a delay of less than 10 ms. In other aspects, the thermal image and the color image can be captured with a delay of more than 1 ns.

[0056] An image pair can be created, for example, by matching the color image to a thermal image taken at the same time. The matching can be achieved, for example, by assigning a first identification to the color image and a second identification to the thermal image taken at the same time as the color image, with the first identification being linked to the second identification.

[0057] According to one aspect, the method comprises: capturing a color calibration image with the first camera of a test object having a plurality of light points; capturing a thermal calibration image with the second camera of the test object having the plurality of light points; calculating, by the computer, the coordinate transformation between a perspective of the first camera and a perspective of the second camera by comparing the positions of the plurality of light points in the color calibration image and the thermal calibration image. The first camera and the second camera can be offset from each other by a small distance. As a result, the perspectives of the two cameras can differ slightly from each other.

[0058] Since the relative positions of the first camera and the second camera are usually fixed, the perspectives can be converted into one another using a coordinate transformation, in particular the affine mapping, for each detection of facial manipulation. The coordinate transformation, in particular the affine mapping, can be determined in a calibration process. For this purpose, a thermal image calibration image and a color image calibration image are each taken of a test object with several spaced-apart light points, with the first and second cameras permanently installed in their future positions. The light points are usually visible in the thermal image and the color image. By superimposing the light points in both images, the coordinate transformation, in particular the affine mapping, can be calculated to align the perspectives of both cameras.

[0059] Thus, the calibration process can be used to determine the coordinate transformation, specifically the affine mapping, which is used to convert the perspectives of both cameras into each other. The calibration process can be repeated for different distances to the cameras and the test object.

[0060] Furthermore, in some aspects, the method may comprise the steps of: selecting the color image if the classification of whether facial manipulation is present results in no facial manipulation; or discarding the color image if the classification of whether facial manipulation is present results in facial manipulation; and generating image data using the selected color image.

[0061] This can provide a method for determining image data representing a digital biometric passport photo for a security document.

[0062] In some aspects, the method further comprises: transmitting the image data to a personalization device configured to personalize a security document, such as an ID card. Such a personalization device can, for example, be designed with a printing device by means of which the color image is printed on the security document. The method can therefore further comprise: printing, by the personalization device, the color image onto a security document based on the image data. Alternatively or additionally, the method can comprise: storing the image data by the personalization device in a storage device of the security document. As a result, the security document is personalized and enables identification of the person based on the image data or the selected color image.

[0063] Such a method and such a personalization device is exemplified in the European application EP 4 099281 A1 , the content of which regarding the personalization device and / or the method for printing a security document and / or the method for determining passport photo data is incorporated herein.

[0064] In other aspects, the method further comprises performing a face recognition process based on the image data.

[0065] Furthermore, the object mentioned at the outset is achieved by a device, in particular a computer, for detecting facial manipulation, which is set up and designed to carry out the method described above, as well as a computer program product comprising instructions which, when executed by a data processing device, in particular the device, cause the latter to carry out the method steps of the method described above.

[0066] Furthermore, the object mentioned at the outset is achieved by a system for detecting facial manipulation, the system comprising: a device as described above; a first camera in connection with the device, the first camera being designed to record light with wavelengths in the visible range and to generate the color image; a second camera in connection with the device, the second camera being designed to record wavelengths in the infrared range and to generate the thermal image; the device being further set up and designed to control the first camera and the second camera such that the first camera records the color image of the face and the second camera records the thermal image of the face at the same time as the color image.

[0067] Such a system can detect facial manipulation. The first camera and the second camera can be aligned in the system such that their fields of view essentially overlap in a plane at a predetermined distance from the two cameras and have a substantially congruent section in the plane.

[0068] In one aspect, the system comprises means arranged and designed to enable a person to independently capture the color image and the thermal image.

[0069] Such a means can be, for example, a trigger connected to the previously described device in such a way that, when the trigger is triggered, a signal is sent to the device, causing a thermal image and a color image to be recorded. The trigger can, for example, be arranged at least in one direction of the fields of view of the first and second cameras, so that a person can independently record a corresponding thermal image and a corresponding color image of their face by triggering the trigger.

[0070] Such a system can also be referred to as a self-enrollment system. This system is specifically designed to allow a person to identify themselves independently without the presence of an authorized person and to detect any tampering. Furthermore, such a system can be the aforementioned personalization device.

[0071] The first camera and the second camera are installed in the system such that a field of view of the first camera and a field of view of the second camera are congruent at a predetermined distance.

[0072] Character list

[0073] In the following, further properties, features and advantages of the invention will become clear by describing preferred embodiments of the invention with reference to the accompanying exemplary drawings, in which: Fig. 1 shows, by way of example, a schematic block diagram for a method for detecting facial manipulation; and

[0074] Fig. 2 shows an exemplary schematic representation of a system for detecting facial manipulation.

[0075] The features disclosed in the above description, the figures and the claims may be important both individually and in any combination for the realization of the invention in the various embodiments.

[0076] Reference symbols in the figures refer to the same elements.

[0077] Figure 1 shows an example of a schematic block diagram of a method for detecting facial manipulation.

[0078] In step S1, a calibration step is performed to calculate a coordinate transformation, in particular an affine transformation such as an affine mapping, between a perspective of the first camera 2, which generates a color image, and a second camera 3, which generates a thermal image. This step is performed in advance for a system with the first and second cameras 2, 3, which are permanently installed, and does not need to be repeated before each detection of facial manipulation to implement the method. Step S1 is therefore optional.

[0079] For calibration, the first camera 2 captures a color calibration image of a test object with a plurality of light points at a predetermined distance from the first camera 2. Furthermore, the second camera 3 captures a thermal image calibration image of the test object with the plurality of light points at the predetermined distance. The light points in both images are superimposed, thus determining a perspective shift between the two captured images. From this, an affine mapping is calculated for a transformation between the perspective of the thermal image and the color image.

[0080] In a step S2, the first camera 2 captures a color image in the visible wavelength range of the face of a person 4 to be identified. Simultaneously with the capture of the color image, the second camera 3 captures a thermal image in the infrared wavelength range of the face of the person 4 to be identified. The face of the person 4 to be identified is located at the predetermined distance from the first and second cameras 2, 3. The color image and the simultaneously captured thermal image are combined so that both images form an image pair.

[0081] In a step S3, specific facial features, such as eyes, nose or mouth, are recognized in the color image using artificial intelligence.

[0082] In a step S4, the color image is scaled to a predetermined distance, for example, on the order of 100 pixels between the eyes of the face in the color image.

[0083] In step S5, the perspective of the thermal image is transformed to a perspective of the color image using coordinate transformation, in particular affine mapping. The thermal image is then scaled to the same dimensions, in particular the same dimensions, as the color image. The thermal image and the color image are thus essentially congruent. In particular, the face in the color image is essentially congruent to the face in the thermal image.

[0084] In step S6, the scaled color image and the scaled thermal image are cropped so that the face in both images is positioned at a predetermined position within the respective image. Cropping the images can be performed, for example, based on ICAO rules. Cropping the face results in a better evaluation. However, the subsequent evaluation can also be performed without cropping. Step S6 is therefore optional.

[0085] In step S7, the scaled color image and the scaled thermal image are combined, for example, stacked. The combined image thus generated preferably essentially corresponds to a tensor that links the information from the color image and the thermal image.

[0086] In a further optional step, the orientation of the head encompassing the face can be determined in the combined image. If the head tilt exceeds a threshold, the corresponding image can be discarded and the process (S2-S7) repeated with another image pair.

[0087] In a step S8, a probability value is determined for each class indicating a type of manipulation by means of a convolved neural network, which indicates the probability that this type of manipulation is present.

[0088] In step S9, the probability values ​​for several classes indicating a type of manipulation are normalized using a softmax function. A normalized probability value indicates a value between 0 and 1.

[0089] In step S10, the normalized probability values ​​for several classes indicating a type of manipulation are summed. If the sum of the normalized probability values ​​exceeds a first threshold, it is classified as facial manipulation. If the sum of the normalized probability values ​​falls below the first threshold or falls below a second threshold that is lower than the first threshold, it is classified as no manipulation.

[0090] Fig. 2 shows an example schematic representation of a system for detecting facial manipulation.

[0091] Such a system comprises a computer 1 which is designed and configured to carry out the method described in connection with Fig. 1.

[0092] A first camera 2 is provided in connection with the computer 1. The first camera 2 is designed to capture a color image in the visible wavelength range. Furthermore, a second camera 3 is provided in connection with the computer 1. The second camera 3 is a thermal imaging camera designed to capture a thermal image in the infrared wavelength range.

[0093] The first camera 1 and the second camera 2 are permanently installed within the system with a fixed position and alignment to one another. The first camera 2 has a first field of view (represented by the corresponding dashed lines), and the second camera 3 has a second field of view (represented by the corresponding dashed lines). The first and second cameras 2, 3 are aligned with one another such that their fields of view are essentially congruent in a plane parallel to the cameras at a predetermined distance from the cameras in a section 5. The face of a person 4 to be identified is positioned in the section 5. Due to the alignment of the first and second cameras 2, 3, the section 5 of the color image of the face of the person 4 to be identified corresponds to the section 5 of the thermal image of the face of the person 4 to be identified.

[0094] Furthermore, the computer 1 is configured to control the first and second cameras 2, 3 such that the color image and the thermal image are recorded simultaneously, i.e., with a negligible time delay. The computer 1 is also configured to identify the simultaneously recorded images and link them together to create an image pair.

[0095] The computer 1 is further designed to carry out the method according to the invention for detecting facial manipulation on the image pair.

[0096] The features disclosed in the above description, the claims and the drawings may be important for the realization of the various embodiments both individually and in any combination.

[0097] List of reference symbols:

[0098] 1 computer

[0099] 2 First camera

[0100] 3 Second Camera 4 Person

[0101] 5 Excerpt

[0102] S1-S10 process steps

Claims

Claims 1. A method for detecting facial manipulation, the method comprising the following steps carried out by means of a computer (1): - Receiving at least one image pair comprising a color image of a face and a simultaneously recorded thermal image of the face; - Adjusting the color image and the thermal image so that the face in the color image is congruent with the face in the thermal image; - Creating a combined image by combining the adjusted color image and the adjusted thermal image; - determining, on the combined image, for one or more classes indicating a type of manipulation, a probability value which indicates a probability that manipulation is present for the corresponding class; and - Classify whether facial manipulation has occurred based on the probability values.

2. The method according to claim 1, wherein adjusting the color image and the thermal image comprises the steps of: - Determining characteristic facial features of the face in the color image; - scaling dimensions of the color image such that a distance between two of the characteristic facial features in the color image has a predetermined value; - perspective transforming the thermal image by means of a coordinate transformation so that a perspective of the thermal image and a perspective of the color image match; and - Scaling dimensions of the thermal image to the dimensions of the scaled color image such that the face in the scaled thermal image is congruent with the face in the scaled color image.

3. The method of claim 2, further comprising: - cropping the scaled color image based on the characteristic facial features so that the face is arranged at a predetermined position and with a predetermined size in the color image; - Cropping the scaled thermal image analogously to cropping the color image, so that the cropped thermal image and the cropped color image have the same section.

4. Method according to one of claims 2 or 3, wherein the characteristic facial features in the color image are determined by means of artificial intelligence.

5. Method according to one of the preceding claims, further comprising: - Normalizing, by the computer, the probability value for each of the one or more classes to a normalized probability value; and - where classifying whether facial manipulation has occurred includes: - Summing the normalized probability values ​​for several classes, - classifying that facial manipulation has occurred if the sum of the normalized probability values ​​is greater than a first threshold, and - Classify that no facial manipulation is present if the sum of the normalized probability values ​​is less than the first threshold or less than a second threshold that is lower than the first threshold.

6. Method according to one of the preceding claims, wherein the determination of a probability value for one or more classes indicating a type of manipulation is carried out by means of a convolutional neural network.

7. Method according to one of the preceding claims, further comprising: - determining, by the computer, on the combined image an orientation of the Frankfurt horizontal of a head encompassing the face relative to a horizontal and an orientation of the frontal plane of the head relative to a vertical; - Rejecting the normalized image if an angle between the Frankfort horizontal of the head and the horizontal exceeds a third threshold and / or if an angle between the frontal plane of the head and the vertical exceeds a fourth threshold.

8. A method according to any one of the preceding claims, wherein the method further comprises: Recording a color image by means of a first camera (2) which is designed to record light with wavelengths in the visible range; and - Recording a thermal image simultaneously with the recording of the color image by means of a second camera (3) which is designed to record wavelengths in the infrared range; - generating, by the computer (1), the at least one image pair by associating the color image and the thermal image recorded at the same time.

9. The method of claim 8, further comprising: - taking a color calibration image with the first camera (2) of a test object having a plurality of light points; - taking a thermal image calibration image with the second camera (3) of the test object with the plurality of light points; - Calculating, by the computer (1), the coordinate transformation between a perspective of the first camera (2) and a perspective of the second camera (3) by comparing the positions of the plurality of light points in the color image calibration image and the thermal image calibration image.

10. Device, in particular a computer (1), for detecting facial manipulation, wherein the device is arranged and designed to carry out the method according to one of claims 1-7.

11. A computer program product comprising instructions which, when executed by a data processing device, in particular the device according to claim 10, cause the device to carry out the method steps according to at least one of claims 1 to 7.

12. A system for detecting facial manipulation, the system comprising: - a device according to claim 10; - a first camera (2) in connection with the device, wherein the first camera (2) is designed to record light with wavelengths in the visible range and to generate the color image; - a second camera (3) in connection with the device, wherein the second camera (3) is designed to record wavelengths in the infrared range and to generate the thermal image; - wherein the device is further configured and designed to control the first camera (2) and the second camera (3) such that the first camera (2) takes a color image of the face and the second camera (3) takes the thermal image of the face at the same time as the color image.

13. The system of claim 12, wherein the system comprises means configured and adapted to enable a person to independently capture the color image and the thermal image.