METHOD AND ARRANGEMENT FOR OPTICALLY DETECTING A PERSON'S HEAD
Patent Information
- Application Number
- DE502022007059
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-04-13
- Filing Date
- 2022-04-12
- Publication Date
- 2026-03-05
- Estimated Expiration
- 2042-04-12
AI Technical Summary
Existing methods for generating three-dimensional facial models for identity documents and access control systems are costly and require complex image processing, often resulting in inaccuracies, especially at the edges of the face, and necessitate the storage of both 3D and 2D images, which can be compromised by occlusions and equipment limitations.
A method involving optical capture of a person's head from multiple directions using a single camera, guided by a visual or auditory stimulus moving along a defined trajectory, generating high-quality three-dimensional models by combining two-dimensional data sets to determine depth information, with quality control for missing or defective pixels.
Enables high-quality three-dimensional model creation with reduced equipment cost and complexity, ensuring accurate capture of facial features even at edges, and allows for efficient storage and comparison with reference data for access control.
Description
[0001] The invention relates to a method for optically detecting a person's head, particularly at an access control station or especially when creating reference data for an identity document, and a corresponding arrangement for carrying out the method.
[0002] Such methods and arrangements are used, for example, to create three-dimensional models or three-dimensional data of a person's head, which can be stored as reference data in an identity document or used for comparison with existing reference data, such as in access control systems, particularly automated access control systems. Often, the portrait image printed on the ID card or personal document is used for comparison. However, this two-dimensional portrait image could be forged, making comparison with other data sensible and practical. In the future, modern facial recognition systems will also use three-dimensional models (3D models) as reference data. For this to work, the acquisition of high-quality reference data is essential.Multi-camera systems are often used to capture 3D models. This allows the front of the head to be viewed simultaneously from various defined angles. Since the positions and viewing directions of each camera are known, the location of every point on the face within the 3D model can be determined. These locations can then be compared with stored reference data. A larger number of cameras generally results in improved accuracy in determining the position of individual facial points. Furthermore, using multiple cameras often allows for the compensation of potential occlusions. However, the disadvantages include the high cost of the equipment and the significant effort required for image processing.
[0003] Furthermore, access control systems (for example, to ensure backward compatibility with existing solutions) often require the storage of a two-dimensional image in addition to the 3D model as reference data. This image is generated by projecting it from the 3D model. The less image information the 3D model contains, and therefore the lower its quality, the more likely inaccuracies will occur, particularly at the edges of the face, such as in the hair or ears. This manifests as dead pixels or even "frayed" ears or hair.
[0004] US patent 2015 / 055085A1 describes a system and method for creating eyeglasses or lenses in which a user's face is captured with a camera, such as a smartphone camera. The user can receive instructions for head movements so that the system can capture the face from as many angles as possible in order to generate and virtually try on suitable eyeglasses.
[0005] In Yuping Lin et al.: "Accurate 3D Face Reconstruction from Weakly Calibrated Wide Baseline Images with Profile Contours" Yuping Lin, Gérard Medioni, and Jongmoo Choi Computer Science Department, University of Southern California 3737 Watt Way, PHE 101, Los Angeles, CA, 90089 978-1-4244-6985-7 / 10 / $26.00 ©2010 IEEE, it is described how a head is optically captured from different directions, and a 3D model of the head can then be created from this captured data.
[0006] It is therefore the object of the present invention to improve a method and an arrangement in such a way that three-dimensional data can be generated with high quality, while requiring little equipment.
[0007] This problem is solved by a method with the features of claim 1 and by an arrangement with the features of claim 8. Advantageous embodiments with expedient further developments of the invention are specified in the dependent claims.
[0008] The procedure includes, in particular, the following steps: Optically capturing a first image of the person's head from a first direction - preferably frontal - and generating a first two-dimensional data set from the first image; moving an optical or acoustic stimulus along a trajectory to a first stop, causing the person to move their head to follow the stimulus to the first stop; optically capturing at least a second image of the person's head from a second direction using the camera or a second camera and generating a second two-dimensional data set from the second image; and determining depth information and / or a three-dimensional model by combining the data sets using an evaluation unit of a processor unit.
[0009] By means of a visual or auditory stimulus that moves along a trajectory to a first stop, it can be ensured that the person's head is moved in a defined manner and can be captured from a desired angle. This enables improved user guidance and ensures that the person's head can be visually captured from as many directions or angles as possible. The stimulus clearly and unambiguously indicates to the person which posture to assume, allowing for the creation of a complete, high-quality three-dimensional model of the head or the acquisition of complete depth information, even at the edges of the face. At the same time, the stimulus allows the process to be carried out with a simplified setup. Thus, a single camera is sufficient for the entire procedure.The acoustic stimulus can be provided, for example, by means of loudspeakers, while the visual stimulus can be output by means of a display element, such as a screen, or by a plurality of light sources, especially LEDs.
[0010] When the procedure is used as part of an access control system, it has proven advantageous to compare the merged data records and the depth information or three-dimensional model with reference values contained or stored on or in an identity document. If the merged data records and the depth information or three-dimensional model deviate from the reference values contained or stored on or in the identity document by a predefined threshold, access is denied. Conversely, access is granted if the merged data records and the depth information or three-dimensional model match the reference values contained or stored on or in the identity document, or if they match each other above a predefined threshold.Furthermore, it is advantageous if the merged datasets and / or the depth information are stored as additional (new) reference data in the access control station's memory. This allows changes in a person's reference data, i.e., their appearance, to be recorded. In the event of future access control, the merged data and depth information, or the resulting three-dimensional model, can then be compared with the additional (new) reference data. This allows access to be granted if there is at least a partial match, or denied if there is a deviation of a predefined threshold.
[0011] When the procedure is used to create reference data for an identity document, it has proven advantageous to store the merged datasets and the depth information or three-dimensional model in or on the identity document, for example, as a biometric three-dimensional image of the person's head. The identity document can then be read using a reader during access control for comparison with the merged datasets and depth information or three-dimensional model.
[0012] Furthermore, it is particularly preferred, if there are several stops along the trajectory, that the optical or acoustic stimulus moves along the trajectory from one stop to the next, thereby prompting the person to move their head in accordance with the stimulus until the next stop. At each stop, another image of the person's head is optically captured by the camera or a second camera, thus generating another two-dimensional data set from the additional image. This enables the optical capture of, in particular, the peripheral areas of the face, such as the chin, neck, ears, or hair.
[0013] To ensure that the person can follow the visual or auditory stimulus unambiguously, it is preferred that the visual or auditory stimulus shifts continuously along its trajectory. In an alternative embodiment, the visual or auditory stimulus can also shift discontinuously along its trajectory, i.e., the stimulus can jump. For example, loudspeakers, screens, display elements, or light sources can be arranged at different locations spaced apart from one another and can be activated or deactivated independently.
[0014] In a particularly simple embodiment, the trajectory is formed by a linear arrangement of light sources or pixels that are activated and / or deactivated in a predetermined sequence. The trajectory can therefore be formed as a band, in particular as an LED band or as a string of lights.
[0015] In an alternative embodiment, the trajectory can be formed as a surface with multiple areas containing light sources or pixels, wherein the areas can be activated and / or deactivated independently. Thus, to initiate head movement, a first area can be activated while the remaining areas remain deactivated, or a first area can be deactivated while the remaining areas remain activated. Alternatively, several, preferably adjacent, areas can be activated while the remaining areas are deactivated. The areas can differ in their emitted intensity or emission spectrum; that is, the areas can emit at a higher intensity in the activated state than in the deactivated state, and vice versa. The areas can be arranged in any two-dimensional configuration, not just along an imaginary straight guide line.This can encourage the user to move their head in a circle, for example.
[0016] To check the quality of the data sets and / or the three-dimensional model, it is preferred that the following further steps be carried out after merging using the evaluation unit: Examining the images for defective pixels and / or missing image information, shifting the optical or acoustic stimulus along the trajectory to another stop or back to a previous stop, causing the person to move their head to follow the stimulus to the next stop or to the previous stop, optically capturing at least one further image using the camera or a second camera and generating another two-dimensional data set from the further image, detecting the missing image information and / or defective pixels from the further data set, and compensating for the detected misinformation and / or the defective pixels in the images.
[0017] This makes it possible to detect defective pixels and / or missing image information, which may be caused by occlusion or by glasses, for example, and to compensate for them with image information from other images / data sets.
[0018] In particular, it is preferred if the procedure includes the following steps: Localization of missing image information and / or defective pixels in the captured images, determination of a head posture to be assumed in which the missing image information and / or defective pixels can be optically detected by the camera, determination of the stop along the trajectory for the stimulus by which the person is caused to assume the posture in which the missing image information and / or defective pixels can be optically detected, and shifting the optical or acoustic stimulus along the trajectory to the determined stop.
[0019] By locating the missing image information in the captured images and by determining the point along the trajectory where the missing image information and / or the defective pixels become visually detectable, it is possible to guide the person to position their head precisely in the posture containing the defective pixels or the incorrect information in order to obtain a complete 3D model or complete depth information. In other words, by locating the missing image information and / or the defective pixels, the head posture that must be assumed to obtain the missing information is predicted.
[0020] According to the invention, a three-dimensional model of the person's head is created from the images, the data sets, and the depth information derived therefrom. Using this three-dimensional model, a two-dimensional image is then generated via projection and stored in a memory. The resulting two-dimensional image exhibits high quality and high resolution, even at the edges.
[0021] The arrangement for carrying out the procedure includes a first camera and a processor unit configured to communicate with the camera via a communication link, the processor unit having an evaluation unit configured to further process data captured from the person by means of the first camera.The first camera is configured to optically capture a first image of the person's head from a first direction – preferably frontally – with which the processor unit generates a first two-dimensional data set from the first image, wherein an attractor is provided which is configured to move an optical or acoustic stimulus along a trajectory to a first stop, thereby causing the person to move their head to follow the stimulus to the first stop, wherein a second image of the person's head can be captured from a second direction by means of the first camera, with which the processor unit generates a second two-dimensional data set from the second image in order to combine the data sets to determine depth information or a three-dimensional model by means of the evaluation unit.
[0022] The device can be part of an access control station, wherein the evaluation unit is configured to process data captured by the first camera of the person being verified by the access control station and subsequently transmit the resulting processing result to an output unit. In addition, a reading unit for reading an identity document is provided, enabling the evaluation unit to compare the data records as well as the depth information or the three-dimensional model with reference values contained on, in, or stored in the identity document.
[0023] The arrangement according to the invention is thus designed such that verification of a person to be checked – preferably at an access control station – can be carried out, whereby a three-dimensional model or a three-dimensional data set can be processed using a small number of cameras. Simple user guidance is possible by means of the attractor, so that the three-dimensional information, i.e., the depth information, can be captured from as many directions as possible. The device can therefore serve both to create a three-dimensional model or a three-dimensional data set that can be used in an identification document, and to identify a person within the context of an access control station.
[0024] In this context, it is preferred if the attractor is formed as a linear arrangement of light sources or pixels, for example, an LED strip or a moving light snake whose illuminated zone moves slowly. Alternatively, the attractor can also be formed in a cross shape, with light sources or pixels arranged in a cross shape.
[0025] In another alternative and preferred embodiment, the attractor can be formed as a surface with a plurality of light sources or pixel-containing areas that can be independently activated or deactivated by the processor unit. The different areas can differ in their emitted intensity or emission spectra. For example, an activated area emits with a higher intensity—i.e., brighter—than the deactivated areas, and vice versa.
[0026] The features and combinations of features mentioned above in the description, as well as those subsequently mentioned in the figure description and / or shown in the figures alone, can be used not only in the combinations specified, but also in other combinations or on their own, without departing from the scope of the invention. Thus, embodiments that are not explicitly shown or explained in the figures, but which can be generated from the explained embodiments, are also to be considered as encompassed and disclosed by the invention.
[0027] Further advantages, features and details of the invention will become apparent from the claims, the following description of preferred embodiments, and the drawings. These show: Fig. 1 a schematic representation of an arrangement for the optical detection of a person to be checked, Fig. 2 a schematic representation of a trajectory with a plurality of stops, Fig. 3 a schematic representation of a linear trajectory with continuous stimulus, Fig. 4 a schematic representation of a linear trajectory with quality control, and Fig. 5 a schematic representation of an area trajectory
[0028] In Figure 1An arrangement 100 for optically capturing the head of a person 200 to be checked at an access control station 104 is shown. The access control station 104 shown here as an example, or the arrangement 100, comprises a first camera 102 and an optional second camera 112, whose optical axes are aligned perpendicular to each other. It should be noted that the use of the second camera 112 is expedient but not strictly necessary, since the head of the person 200 to be checked could also be captured multiple times by the first camera 102.Both cameras 102, 112 are connected via a communication link shown in dotted lines to a processor unit 106, which in turn has an evaluation unit 108 to further process the data that is captured by the camera 102, 112 of the person 200 to be checked, so that it can then be incorporated into a processing result and, after transmission to an output unit 110, output by the latter.
[0029] The first camera 102 is configured to capture a first image of the person 200 to be checked, this first image being preferably a frontally oriented portrait image of the person 200. The processor unit 106 then generates a first two-dimensional data set from the first image. The access control station 104, or the arrangement 100, also includes an attractor 300 configured to move an optical or acoustic stimulus along a trajectory 301 to a first stop 303, causing the person 200 to move their head to follow the stimulus to the first stop 303. The first camera 102 or the second camera 112 is configured to capture a second image of the head of the person 200 to be checked from a second direction – preferably from a profile.The processor unit 106 generates a second two-dimensional data set from the second image, in order to subsequently combine the data sets to determine depth information or to generate a three-dimensional model using the evaluation unit 108.
[0030] The newly acquired data set can then be compared with information on an identity document, in particular information stored on a chip of the identity document. For this comparison, the arrangement 100 additionally includes a reading unit (not shown) for reading the identity document, which enables the evaluation unit 108 to compare the data sets as well as the depth information or the 3D model with reference values contained on or in the identity document or stored therein, in order to output the comparison result to the output unit 110 and subsequently grant or deny access.
[0031] In order to capture the head of the person 200 being checked from as many angles and directions as possible, and to improve the quality of the three-dimensional information, the attractor 300 has a trajectory 301 with several stops 302, so that the optical or acoustic stimulus moves along the trajectory 301 from one of the stops 302 to the next of the stops 304. This causes the person 200 to move their head in response to the stimulus until they reach the next stop 304, whereby at each of the stops 302 another image of the head of the person 200 being checked can be optically captured by the first camera 102 or the second camera 112, thus generating another two-dimensional data set from the additional image.
[0032] Fig. 3Figure 1 shows an attractor formed as a linear arrangement 306 of light sources 305 or pixels. For example, the attractor 300 can be formed as an LED strip or as a string of lights. Figure 3 A temporal progression of the optical stimulus along trajectory 301 is depicted. The optical stimulus shifts continuously along trajectory 301; that is, the light sources 305 are deactivated along trajectory 301 in a predefined sequence, i.e., one after the other. Activation, or turning on, refers to the emission of the light source 305 at a higher intensity, while deactivation of the light sources 305 means that they do not emit at all or at least emit at a lower intensity. Alternatively, the light sources 305 or pixels can, of course, also be switched sequentially from a deactivated state to an activated state.
[0033] In the Figure 4The attractor 300 is again represented as a linear arrangement 306, where a temporal progression of the optical stimulus is also depicted. In this case, the stimulus jumps along the trajectory 301, meaning the progression is discontinuous. This is particularly advantageous when the process needs to be accelerated and when quality control is performed. For quality control of the generated depth information or the generated three-dimensional model, the evaluation unit 108 is configured to examine the images and / or the data sets and / or the three-dimensional model for defective pixels and / or missing image information. Missing image information can also include, for example, occlusions caused by eyeglasses or hair.If defective pixels or missing image information are present, the processor unit 106 instructs the attractor 300 to move the optical or acoustic stimulus along the trajectory 301 to a further stop 302 or back to a previous stop 302. In this case, the optical stimulus was moved back to a previous stop 302, causing the person 200 to move their head in response to the stimulus. Subsequently, further images can be optically captured using the first camera 102 or the second camera 112 to generate additional two-dimensional datasets, or the optical capture of images from a specific direction can be repeated. From the additional datasets thus generated, the missing image information and / or the defective pixels can be detected and compensated for in the depth information or in the three-dimensional model.
[0034] Preferably, the missing image information in the captured images is located by means of the evaluation unit 108, which is configured to determine a head posture in which the missing image information and / or the defective pixels can be optically detected by camera 102 or the second camera 112. The evaluation unit 108 also determines the stop 302 along the stimulus trajectory that causes person 200 to assume the posture in which the missing image information and / or the defective pixels can be optically detected. Subsequently, the optical or acoustic stimulus is moved along the trajectory 301 to the determined stop 302, and further images are optically captured.
[0035] Of course, the displacement of the optical or acoustic stimulus 301 along the trajectory 301 can also be carried out continuously as part of quality control.
[0036] The resulting high-quality 3D depth information or 3D model of the person's head can then be converted into a two-dimensional image by projection and stored in a memory.
[0037] Figure 5 Figure 1 shows another embodiment of the attractor 300, which in this case is formed as a surface 307, with areas 308 having a plurality of light sources 305. The areas 308 can be activated and / or deactivated independently of one another, i.e., they can be switched from bright to dark or from higher intensity to lower intensity. REFERENCE MARK LIST
[0038] 100 Arrangement 102 First camera 104 Access control station 106 Processor unit 108 Evaluation unit 110 Output unit 112 Second camera 200 Person to be checked 300 Attractor 301 Trajectory 302 Stop 303 First stop 304 Next stop 305 Light source 306 Linear arrangement 307 Area 308 Regions
Claims
1. A method for optically capturing a persons' head (200) as part of access control or for creating reference data for an identity document, comprising the following steps: - optically capturing a first image of the persons' head (200) from a first direction using a camera (102) and generating a first two-dimensional data set from the first image, - moving an optical or acoustic stimulus along a trajectory (301) to a first stop (303), thereby causing the person (200) to move their head to follow the stimulus to the first stop (303), - optically capturing at least a second image of the persons' head (200) from a second direction using the camera (102) or using a second camera (112) and generating a second two-dimensional data set from the second image, and - determining a three-dimensional model of the persons' head by merging the first two-dimensional data set and the second two-dimensional data set by means of an evaluation unit (108) of a processor unit (106) based on the previously known viewing direction and position of the camera (102) or cameras (102; 112), wherein a two-dimensional image is generated by means of projection using the three-dimensional model of the persons' head and stored in a memory.
2. The method according to claim 1, characterized in that there are several stops (302) along the trajectory (301), that the optical or acoustic stimulus moves along the trajectory (301) from one of the stops (302) to the next of the stops (304), causing the person (200) to a movement of the head in order to follow the stimulus to a next of the stops (304), and that at each of the stops (302), another image of the head of the person being examined (300) is optically captured by means of the camera (102) or a second camera (112), thereby generating a further two-dimensional data set from the further image.
3. The method according to claim 1 or 2, characterized in that the optical or acoustic stimulus moves continuously along the trajectory (301).
4. The method according to any one of claims 1 to 3, characterized in that the trajectory (301) is formed by a linear arrangement (306) of light sources (305) or pixels, which are activated and / or deactivated in a predetermined sequence.
5. The method according to any one of claims 1 to 3, characterized in that the trajectory (301) is formed as a surface (307) with a plurality of areas (308) comprising light sources (305) or pixels, and in that the areas (308) are activated and / or deactivated independently of one another.
6. The method according to any one of claims 1 to 5, characterized in that, following the merging by means of the evaluation unit (108), the following further steps are carried out: - examining the images for defective pixels and / or missing image information, - moving the optical or acoustic stimulus along the trajectory (301) to another stop (302) or back to a previous one of the stops (302), thereby causing the person (200) to move their head to follow the stimulus to the further stop (302) or to the previous stop (302), - optically capturing at least one further image using the camera (102) or using a second camera (112) and generating a further two-dimensional data set from the further image, - detecting the missing image information and / or missing pixels from the further data set, and - compensating for the detected missing information and / or missing pixels in the images.
7. The method according to any one of claims 1 to 6, characterized by the following steps: - localizing missing image information and / or defective pixels in the captured images, - determining a position to be assumed by the head in which the missing image information and / or the defective pixels can be optically captured by the camera (102) or a second camera (112), - determining the stop (302) along the trajectory (301) for the stimulus at which the person (200) is prompted to adopt the position in which the missing image information and / or the defective pixels can be optically captured, and - moving the optical or acoustic stimulus along the trajectory (301) to the determined stop (302).
8. An arrangement (100) for carrying out the method according to any one of claims 1 to 7, with a first camera (102) and with a processor unit (106) which is designed to communicate with the camera (102) via a communication link, wherein the processor unit (106) has an evaluation unit (108) which is designed to further process data captured by the first camera (102) from the person (200) to be checked, wherein the first camera (102) is designed to optically capture a first image (114) of the head of the person (200) from a first direction, whereby the processor unit (106) generates a first two-dimensional data set from the first image (114), wherein an attractor (300) is provided which is arranged to move an optical or acoustic stimulus along a trajectory (301) to a first stop (303), causing the person (200) to a movement of the head in order to follow the stimulus to the first stop (303), wherein a second image (116) of the head of the person (200) can be captured from a second direction by means of the first camera (102), whereby the processor unit (106) generates a second two-dimensional data set from the second image (116) in order to combine the first two-dimensional data set and the second two-dimensional data set to determine a three-dimensional model by means of the evaluation unit (108), wherein a two-dimensional image is generated by means of the three-dimensional model of the persons' head by means of projection and stored in a memory.
9. The arrangement (100) according to claim 8, characterized in that the attractor (300) is formed as a linear arrangement (306) of light sources (305) or pixels.
10. The arrangement (100) according to claim 8, characterized in that the attractor (300) is formed as a surface (307) with a plurality of areas (308) comprising light sources (305) or pixels, which can be activated or deactivated independently of one another by means of the processor unit (106).