DEVICE, METHOD AND COMPUTER PROGRAM FOR CORRECTING A PERSON'S FACIAL IMAGE
Patent Information
- Application Number
- DE502018016277
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2017-06-29
- Filing Date
- 2018-06-27
- Publication Date
- 2026-01-08
- Estimated Expiration
- 2038-06-27
AI Technical Summary
Existing facial image capture technologies face challenges in ensuring standardized and high-quality images for automated facial recognition, particularly in limited spaces and with non-standardized distances, leading to higher classification errors and reduced user-friendliness and throughput in security-critical applications.
A device and method that corrects facial images by performing a three-dimensional image transformation to align the captured image with a predefined spatial position, using a spatial orientation determination device to determine the initial distance and orientation, and applying perspective distortion to achieve a standardized facial image.
Enables the capture of facial images from any distance and perspective, ensuring high-quality, standardized images suitable for automated facial recognition systems, reducing false acceptance and rejection rates, and enhancing user-friendliness and throughput.
Description
[0001] The present invention relates to the field of capturing facial images for identification documents, in particular identity cards and passports.
[0002] The patent application EP 2 755 164 A2 concerns the field of video communication.
[0003] The patent application US 2016 / 335495 A1 discloses a device and a method for capturing a facial image.
[0004] The publication "PHILIPP FECHTELER ET AL: "Fast and High Resolution 3D Face Scanning", IMAGE PROCESSING, 2007. ICIP 2007. IEEE INTERNATIONAL CONFERENCE ON, IEEE, Pl, September 1, 2007 (2007-09-01), pages III-81, XP031158009, ISBN: 978-1-4244-1436-9)" deals with the acquisition of 3D facial models.
[0005] Disclosure document US 2007 / 104362 A1 discloses a facial recognition method.
[0006] For various applications in the field of personal identification, the unambiguous assignment of individuals to their respective identification documents is of particular importance. Standardized facial images are especially desirable as reference data for automated facial recognition. Various standardization bodies dealing with the standardization of facial images for identification documents, such as ISO and ICAO, are currently discussing standardized distances and perspectives for capturing facial images. Among other things, the question arises whether a minimum distance between a camera and a person's face should be mandatory for capturing the facial image.
[0007] Furthermore, it must be considered that insufficient image quality typically leads to higher classification errors in automated facial recognition. Especially for security-critical applications, such as border control, legislators prescribe maximum false acceptance rates, for example, 0.1%. Additionally, unsuitable facial images typically result in higher false rejection rates. This negatively impacts both the user-friendliness of automated facial recognition systems and their throughput, as it may necessitate manual follow-up checks.
[0008] When capturing facial images for identification documents, it should also be considered that the available space in photo kiosks or photo booths is usually limited. Furthermore, capturing facial images for non-governmental applications using smartphones, where individuals can take selfies with a built-in camera, is desirable.
[0009] It is therefore an object of the present invention to create an efficient concept for capturing and correcting a person's facial image.
[0010] This task is solved by the features of the independent claims. Advantageous further developments are the subject of the dependent patent claims, the description, and the drawings.
[0011] The invention is based on the finding that the above problem can be solved by perspective distortion of a captured facial image of a person, whereby a three-dimensional image transformation can be performed. The facial image can be captured in any first spatial position between the camera and the person's face and perspectively distorted such that the corrected facial image is assigned to a second spatial position. The second spatial position can, for example, be predefined by a standard. For determining the first spatial position, for example, a distance between the camera and the person's face, as well as a facial shape of the person's face, can be determined, which can be compared with a reference facial shape. This allows, among other things, the correction of the first spatial position.The fact that the facial image corresponds to a projection of the person's three-dimensional face onto a two-dimensional image plane is exploited. The perspective distortion of the facial image can be achieved, for example, by means of a three-dimensional image transformation or coordinate transformation.
[0012] According to a first aspect, the invention relates to a device for correcting a person's facial image according to claim 1.
[0013] According to one embodiment, the spatial orientation determination device is arranged in a fixed position relative to the camera. This achieves the advantage that an auxiliary spatial orientation of the spatial orientation determination device relative to the person's face can first be determined, and the first spatial orientation of the camera relative to the person's face can be determined using the previously known fixed arrangement and the auxiliary spatial orientation.
[0014] The spatial orientation determination device is designed to determine the initial distance between the camera and the person's face, and to further determine the initial spatial orientation using this distance. This achieves the advantage of efficiently scaling the facial image to compensate for perspective distortion.
[0015] The spatial orientation device is designed to detect a person's pair of eyes in the facial image, determine the distance between those eyes in the facial image, and calculate the initial distance between the camera and the person's face based on this distance. This achieves the advantage of efficient distance determination.
[0016] The spatial orientation device can compare the interpupillary distance in the facial image with a predetermined reference interpupillary distance, and thus determine the distance. The reference interpupillary distance can, for example, be a value between 60 mm and 65 mm.
[0017] According to one example, the spatial orientation device is configured to receive an image sharpness indicator from the camera, where the image sharpness indicator represents the sharpness of the face image, and to determine the distance between the camera and the person's face based on this indicator. This achieves the advantage that information about the camera's focus area or focal plane can be evaluated to determine the distance.
[0018] For example, the spatial orientation device is configured to emit a light beam towards the person's face, receive a reflected light beam from the person's face, and determine the distance between the camera and the person's face based on the light beam and the reflected light beam. This achieves the advantage of efficient distance determination.
[0019] The spatial orientation determination device can include a distometer or be a distometer in itself. The light beam and the reflected light beam can be a laser light beam and a reflected laser light beam, respectively.
[0020] The spatial orientation determination device is designed to determine a three-dimensional facial shape of the person's face and to further determine the initial spatial orientation using this three-dimensional facial shape. This achieves the advantage of efficiently determining the initial spatial orientation.
[0021] The spatial orientation determination device is designed to compare the three-dimensional facial shape of the person's face with a predetermined three-dimensional reference facial shape, and to determine the initial spatial orientation based on this comparison. This achieves the advantage of efficiently determining the initial spatial orientation.
[0022] The spatial orientation device is designed to compare the characteristics of the three-dimensional facial shape with the three-dimensional reference facial shape. This can include, for example, the position of the eyes, nose, or ears, and / or the shape of the cheeks, temples, or forehead.
[0023] According to one example, the device comprises a further image camera, wherein the further image camera is configured to capture another image of the person's face in order to obtain a further facial image, and wherein the spatial orientation determination device is configured to determine the shape of the person's face based on the facial image and the further facial image. This achieves the advantage that the facial shape can be determined efficiently.
[0024] The spatial orientation device can include a stereo camera or be a stereo camera in its own right. The main camera and the additional camera can be fixed in place. The spatial orientation device can take into account the different perspectives of the main camera and the additional camera when capturing the facial image or the additional facial image relative to the person's face.
[0025] For example, the spatial orientation determination device is configured to emit a light pattern towards the person's face, receive a reflected light pattern from the person's face, and determine the shape of the person's face based on the emitted and reflected light patterns. This achieves the advantage of efficiently determining the face shape.
[0026] The spatial orientation device can include or be a time-of-flight (TOF) camera. The light pattern can have a predetermined spatial structure, and the spatial orientation device can compare the spatial structure of the reflected light pattern with the predetermined spatial structure of the emitted light pattern. In particular, light travel times or differences in light travel times can be detected and processed.
[0027] According to one example, the light pattern and the reflected light pattern comprise light from a predetermined wavelength range, in particular from a near-infrared (NIR) wavelength range. This achieves the advantage that invisible light can be used to determine facial shape. The light pattern and the reflected light pattern can exclusively comprise light from the predetermined wavelength range, in particular the near-infrared (NIR) wavelength range.
[0028] According to one embodiment, the processor is configured to determine a person's characteristic, in particular their age, gender, mass, or body mass index, based on their facial image, and to select the predetermined reference face model from a plurality of predetermined reference face models based on the person's characteristic. This achieves the advantage that the reference face model used can be selected specifically for each individual.
[0029] According to one embodiment, the processor is configured to capture the person's face in the corrected facial image, determine a background image within the corrected facial image (where the background image does not include the person's face), and remove the background image from the corrected facial image. This achieves the advantage that the corrected facial image can be provided in a standardized manner, i.e., regardless of the background present when the facial image is captured.
[0030] According to one embodiment, the device comprises a communication interface configured to establish a communication connection to a server via a communication network and to transmit the corrected facial image to the server via this communication connection. This achieves the advantage that a large number of facial images from a large number of people can be centrally stored and further processed by the server.
[0031] The communication connection can be an authenticated communication connection. During the establishment of the communication connection, the device can authenticate itself to the server. Furthermore, the server can authenticate itself to the device during the establishment of the communication connection. The authenticated communication connection can be secured using transport encryption.
[0032] The first and second spatial positions are defined by at least one of the following spatial position parameters: a first or second distance between the camera and the person's face, a first or second azimuth angle of the camera relative to the person's face, a first or second elevation angle of the camera relative to the person's face, a first or second orientation of the camera, or a first or second orientation of the person's face. This achieves the advantage that the respective distance and perspective relative to the person's face can be unambiguously defined.
[0033] According to one embodiment, the processor is configured to provide the corrected facial image in accordance with the ISO / IEC 19794-5 standard or the ISO / IEC 39794-5 standard. This offers the advantage that the corrected facial image can be used directly for an identification document.
[0034] According to one embodiment, the device is a smartphone, a photo kiosk, or a self-service terminal. This offers the advantage of being able to utilize established or widely used architectures for implementing the device.
[0035] According to a second aspect, the invention relates to a method for correcting a person's facial image according to claim 4.
[0036] The method can be carried out by the device defined in claim 4. Further features of the method result directly from the features and / or functionality of the device.
[0037] According to a third aspect, the invention relates to a computer program with program code for executing the method according to the second aspect of the invention. The device can be programmed to execute the computer program. The computer program can, for example, be implemented in the form of an app on a smartphone.
[0038] The invention can be implemented in hardware and / or software.
[0039] Further examples of implementation are explained in more detail with reference to the accompanying drawings. These show: Fig. 1 a schematic diagram of a device for correcting a person's facial image; Fig. 2 a schematic diagram of a procedure for correcting a person's facial image; and Fig. 3 A schematic diagram of a facial image correction.
[0040] Fig. 1 Figure 1 shows a schematic diagram of a device 100 for correcting a person's facial image. The device 100 comprises an image camera 101 configured to capture an image of the person's face in order to obtain the facial image, wherein the image camera 101 is arranged in a first spatial orientation relative to the person's face. The device 100 further comprises a spatial orientation determination device 103 configured to determine the first spatial orientation using a predetermined reference face model, wherein the predetermined reference face model represents a predetermined reference face shape.The device 100 also includes a processor 105, which is configured to determine a spatial orientation deviation of a second spatial orientation from the first spatial orientation and to perspectively distort the facial image based on the spatial orientation deviation between the second and first spatial orientations in order to obtain a corrected facial image of the person. The device 100 may further include a communication interface 107, which is configured to establish a communication connection to a server via a communication network and to transmit the corrected facial image to the server via the communication connection.
[0041] Fig. 2 Figure 200 shows a schematic diagram of a method for correcting a person's facial image using a device comprising a camera, a spatial orientation device, and a processor. The camera is arranged in a first spatial orientation relative to the person's face.The procedure 200 comprises capturing 201 an image of the person's face by the camera to obtain the facial image, determining 203 the first spatial orientation using a predetermined reference face model by the spatial orientation determination device, wherein the predetermined reference face model represents a predetermined reference face shape, determining 205 a spatial orientation deviation of a second spatial orientation from the first spatial orientation by the processor, and perspectivally distorting 207 the facial image on the basis of the spatial orientation deviation between the second spatial orientation and the first spatial orientation by the processor to obtain a corrected facial image of the person.
[0042] Fig. 3Figure 301 shows a schematic diagram of the correction of a facial image 301, which was acquired from an initial spatial orientation of an image camera relative to a person's face. The initial spatial orientation is defined by the initial distance between the image camera and the person's face, the initial azimuth angle of the image camera relative to the person's face, the initial elevation angle of the image camera relative to the person's face, the initial orientation of the image camera, and the initial orientation of the person's face. The initial spatial orientation is determined by a spatial orientation determination device, which can, for example, determine the shape of the face and compare it with a predetermined reference face shape.
[0043] Subsequently, a spatial orientation deviation between a second spatial orientation and the first spatial orientation is determined, and the facial image 301 is perspectively distorted by a processor such that a corrected facial image 303 is provided, which corresponds to the second spatial orientation. The second spatial orientation is defined by a second distance between the camera and the person's face, a second azimuth angle of the camera relative to the person's face, a second elevation angle of the camera relative to the person's face, a second orientation of the camera, and a second orientation of the person's face. The second spatial orientation can be predefined and, for example, taken from the ISO / IEC 19794-5 standard.
[0044] The concept therefore allows the provision of a perspectively or geometrically corrected facial image 303, which can be used, for example, as a reference for applications in the field of facial biometrics. The person's face can thus be captured from any distance, for example, an arm's length, and from any perspective to obtain the facial image 301. The captured facial image 301 can then be corrected or harmonized in such a way that it appears as if the face was captured from a freely selectable, optimal distance, for example, between 1.2 m and 2.5 m, preferably 1.5 m, and from a freely selectable, optimal perspective, without impairing the resolution of the facial details.The correction may reduce the resolution of the corrected facial image 303; however, this reduction can be compensated for in advance by an increased resolution when capturing the facial image 301.
[0045] The reduction in resolution caused by perspective distortion can therefore be compensated for by using a higher-resolution optical camera from the outset than would typically be used for an unmodified photograph. For example, the ISO / IEC 39794-5 standard requires 90 pixels between the person's eyes and recommends 120 pixels, ensuring that facial structures remain clearly recognizable, provided they are at least 1 mm in size. Particular care should be taken to ensure that small details of the person's face are not lost. If, for instance, a pixel contains information about a skin fold and this pixel is removed during perspective distortion, the skin fold may disappear.It should therefore be ensured that the initial resolution of the face image 301 is higher than the final required resolution of the corrected face image 303, for example by a factor of between 2 and 3. In order to be able to carry out the perspective distortion of the face image 301, it is advantageous to know the actual spatial structure or shape of the face and the distance of the face from the camera as precisely as possible.
[0046] For example, it is possible to estimate the distance based on the spatial conditions in a photo booth or, in the case of self-portraits, on the average typical arm length of a person. The spatial structure or facial shape can be estimated, for example, using a predetermined, generic facial model.
[0047] The distance between the person's face and the camera can also be estimated based on the interpupillary distance (IPD) in the facial image 301, which can be compared to a reference IPD, for example, between 60 mm and 65 mm. Furthermore, the distance can be determined using a distometer or a laser rangefinder, although care must be taken to ensure that the light or laser light is not directed into the person's eye. For example, the light or laser light can be directed at the person's forehead. The distance can also be determined using the camera's focus indicator or its focus setting when capturing the facial image 301. Finally, the distance can be determined using a time-of-flight (TOF) camera, which can simultaneously determine the shape of the face.Furthermore, the distance can be determined using two facial images captured from different distances. For this purpose, for example, the image capture device according to DE 10 2015 106 358 A1 can be used, whereby the facial shape can be determined simultaneously.
[0048] The geometry or shape of a person's face can be compared or approximated using a number of reference face models or prototypes. A suitable reference face model can be selected based on factors such as age, gender, ethnicity, and / or body mass index. Established facial recognition techniques can be used to determine these characteristics. As previously explained, facial shape can also be determined using a time-of-flight (TOF) camera. In particular, facial shape can be determined by emitting or projecting and receiving structured light, for example, from the near-infrared (NIR) wavelength range. Furthermore, facial shape can be determined by simultaneously capturing two facial images from different distances.
[0049] Knowing the distance and the face shape or geometry, the pixels of the originally captured face image 301 can be rearranged so that their position in the corrected face image 303 corresponds to viewing the face in the optimal, freely selectable second spatial orientation. For example, the person's nose can be made smaller, the surrounding areas slightly enlarged, and the facial edges, cheeks, and ears significantly enlarged. The perspective distortion of the face image 301 can be performed, for example, by means of a three-dimensional image transformation or coordinate transformation. Consequently, the three-dimensional face shape and a three-dimensional reference face shape can be used to determine the first spatial orientation.
[0050] Information about a person's ears, which may be obscured by other parts of the face at close range, could potentially be lost due to perspective distortion. However, typical facial recognition approaches primarily use reference data from a rectangular region of the face bounded by the person's eyes and mouth. Consequently, the performance of typical facial recognition approaches can be maintained.
[0051] The concept thus enables the provision of a distance-harmonized, corrected facial image 303, whereby the facial image 301 can be captured within a wide range of possible, and especially short, distances. The concept can be used, for example, for capturing a facial image 301 in a self-service terminal or a photo kiosk. Different, even very short, distances during the capture of the facial image 301 can be compensated for.
[0052] In an example application at a self-service terminal for passport applications in a residents' registration office or citizens' service center, it is known in advance that the person is approximately an arm's length away from the camera. A precise distance can be determined, for example, by a combination of measuring the interpupillary distance in the facial image (301) and / or based on an image sharpness indicator or distance data from the focused camera. The facial shape or geometry can be determined using structured NIR light. To receive the NIR light, a suitable narrowband filter can be temporarily placed in front of the camera, for example, by folding it down.Using the specific three-dimensional facial shape, the captured facial image 301 can then be harmonized in high resolution to an apparent distance of 150cm and a desired perspective, for example a frontal perspective.
[0053] All features described or shown in connection with individual embodiments can be provided in any combination in the object according to the invention in order to simultaneously realize their advantageous effects. REFERENCE MARK LIST
[0054] 100 Device 101 Camera 103 Spatial orientation device 105 Processor 107 Communication interface 200 Methods for correcting a person's facial image 201 Capturing an image of a person's face 203 Determining an initial spatial orientation 205 Determining a spatial orientation deviation 207 Perspective distortion 301 Facial image 303 Corrected facial image
Claims
1. A device (100) for correcting a facial image (301) of a person, comprising: an image camera (101), which is configured to capture an image of a face of the person in order to obtain the facial image (301), wherein the image camera (101) is arranged in a first spatial position relative to the face of the person; a spatial position determination device (103), which is configured to determine the first spatial position using a predetermined reference face model, wherein the predetermined reference face model represents a predetermined three-dimensional reference face shape; and a processor (105), which is configured to determine a spatial position deviation of a second spatial position from the first spatial position, and to perspectively distort the facial image (301) based on the spatial position deviation between the second spatial position and the first spatial position in order to obtain a corrected facial image (303) of the person, wherein the perspective distortion of the facial image is carried out by means of a three-dimensional image transformation, where the corrected facial image (303) is assigned to the second spatial position, where the second spatial position is predetermined, where the first spatial position and the second spatial position are different, characterized in that the first spatial position is defined by the following spatial position parameters: a first distance between the image camera (101) and the face of the person, a first azimuth angle of the image camera (101) relative to the face of the person, a first elevation angle of the image camera (101) relative to the face of the person, a first orientation of the image camera (101), a first orientation of the face of the person, such that the first distance and the perspective corresponding to the first azimuth angle of the image camera (101), the first elevation angle of the image camera (101), the first orientation of the image camera (101) and the first orientation of the face of the person are uniquely determined, wherein the second spatial position is defined by the following spatial position parameters: a second distance between the image camera (101) and the face of the person, a second azimuth angle of the image camera (101) relative to the face of the person, a second elevation angle of the image camera (101) relative to the face of the person, a second orientation of the image camera (101), a second orientation of the face of the person, such that the second distance and the perspective corresponding to the second azimuth angle of the image camera (101), the second elevation angle of the image camera (101), the second orientation of the image camera (101) and the second orientation of the face of the person are uniquely determined, wherein the spatial position determination device (103) is configured to determine the first distance between the image camera (101) and the face of the person, and to determine the first spatial position further by using the first distance between the image camera (101) and the face of the person, wherein the spatial position determination device (103) is configured to detect a pair of eyes of the person in the facial image (301), to determine an eye-pair distance of the pair of eyes of the person in the facial image (301), and to determine the first distance between the image camera (101) and the face of the person on the basis of the eye-pair distance in the facial image (301), wherein the spatial position determination device (103) is configured to determine a three-dimensional facial shape of the face of the person, and to determine the first spatial position further by using the three-dimensional facial shape of the face of the person, wherein the spatial position determination device is configured to compare the three-dimensional facial shape of the face of the person with the predetermined three-dimensional reference facial shape, and to determine the first spatial position on the basis of comparing the three-dimensional facial shape of the face of the person with the predetermined three-dimensional reference facial shape, wherein for this, the spatial position determination device is configured to compare characteristics of the three-dimensional facial shape and the three-dimensional reference facial shape.
2. The device (100) according to claim 1, wherein the processor (105) is configured to determine a characteristic of the person, in particular an age of the person, a gender of the person, a mass of the person or a body mass index of the person, on the basis of the facial image (301) of the person, and to select the predetermined reference face model from a plurality of predetermined reference face models on the basis of the characteristic of the person.
3. The device (100) according to one of the preceding claims, comprising: a communication interface (107), which is configured to establish a communication connection to a server via a communication network and to transmit the corrected facial image (303) to the server via the communication connection.
4. A method (200) for correcting a facial image (301) of a person using a device (100), wherein the device (100) comprises an image camera (101), a spatial position determination device (103), and a processor (105), wherein the image camera (101) is arranged in a first spatial position relative to a face of the person, the method (200) comprising: capturing (201) an image of the face of the person by the image camera (101) in order to obtain the facial image (301); determining (203) the first spatial position using a predetermined reference face model by the spatial position determination device (103), wherein the predetermined reference face model represents a predetermined three-dimensional reference face shape; determining (205) a spatial position deviation of a second spatial position from the first spatial position by the processor (105); and perspectively distorting (207) the facial image (301) based on the spatial position deviation between the second spatial position and the first spatial position by the processor (105) in order to obtain a corrected facial image (303) of the person, wherein the perspectively distorting of the facial image is carried out by means of a three-dimensional image transformation, where the corrected facial image (303) is assigned to the second spatial position, where the second spatial position is predetermined, where the first spatial position and the second spatial position are different, characterized in that the first spatial position is defined by the following spatial position parameters: a first distance between the image camera (101) and the face of the person, a first azimuth angle of the image camera (101) relative to the face of the person, a first elevation angle of the image camera (101) relative to the face of the person, a first orientation of the image camera (101), a first orientation of the face of the person, such that the first distance and the perspective corresponding to the first azimuth angle of the image camera (101), the first elevation angle of the image camera (101), the first orientation of the image camera (101) and the first orientation of the face of the person are uniquely determined, wherein the second spatial position is defined by the following spatial position parameters: a second distance between the image camera (101) and the face of the person, a second azimuth angle of the image camera (101) relative to the face of the person, a second elevation angle of the image camera (101) relative to the face of the person, a second orientation of the image camera (101), a second orientation of the face of the person, such that the second distance and the perspective corresponding to the second azimuth angle of the image camera (101), the second elevation angle of the image camera (101), the second orientation of the image camera (101) and the second orientation of the face of the person are uniquely determined, wherein, by the spatial position determination device (103), the first distance between the image camera (101) and the face of the person is determined and the first spatial position is further determined using the first distance between the image camera (101) and the face of the person, wherein, by the spatial position determination device (103), a pair of eyes of the person in the facial image (301) is detected, an eye-pair distance of the pair of eyes of the person in the facial image (301) is determined, and the first distance between the image camera (101) and the face of the person based on the eye-pair distance is determined in the facial image (301), wherein, by the spatial position determination device (103), a three-dimensional facial shape of the face of the person is determined and the first spatial position is further determined using the three-dimensional facial shape of the face of the person, wherein, by the spatial position determination device, the three-dimensional facial shape of the face of the person is compared with the predetermined three-dimensional reference facial shape and the first spatial position is determined based on the comparison of the three-dimensional facial shape of the face of the person with the predetermined three-dimensional reference facial shape, wherein, by the spatial position determination device, characteristics of the three-dimensional facial shape and the three-dimensional reference facial shape are compared for this.
5. A computer program comprising program code for executing the method (200) according to claim 4, when the method (200) is executed by the device (100) according to any one of claims 1 to 3, wherein the device (100) is programmably configured to execute the computer program.