Improved liveness detection

The liveness detection method uses a flash setting sequence to detect eye reflections and apply filters for geometric analysis, effectively distinguishing real from fake faces, enhancing the resistance of facial biometric systems to spoofing attacks.

JP2026070491APending Publication Date: 2026-04-27AMADEUS SAS +1
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
AMADEUS SAS
Filing Date
2025-10-14
Publication Date
2026-04-27

Smart Images

  • Figure 2026070491000001_ABST
    Figure 2026070491000001_ABST
Patent Text Reader

Abstract

This invention provides a liveness detection method for determining whether a face presented to a user is a real face. [Solution] The method obtains each image in an image sequence, including left eye image data and right eye image data, from a flash setting sequence, from a face presented under flash settings including different flash setting values. A filter is applied to suppress non-spotted image data compared to spotted image data in order to enhance spotted image data corresponding to spots that are brighter than the flash setting for subsequent images, resulting from the eyes reflecting the flash. When one left eye spot and one right eye spot are detected from the left eye image data and right eye image data, respectively, a liveness detection result is output based on a determination of whether the detected left eye spot and right eye spot match each other.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an automatic face authentication or identification system, particularly for addressing potential vulnerabilities of an automatic face authentication or identification system to spoofing attacks.

Background Art

[0002] The use of biometrics in personal authentication or identification has been increasing in recent years, particularly against the backdrop of advancements in face recognition and image processing technologies. Applications where such use can be easily adopted include passenger identification or registration at airports, especially where there are already self-service kiosk terminals where passengers can complete other functions such as checking in for a flight, printing boarding passes, or printing baggage tags. With the progress of computing and camera technologies, face biometrics verification is also likely to be increasingly used in other scenarios such as building access control.

[0003] In a face biometrics identification system, an image of a person's face is captured, analyzed, and compared with a database of registered face data to determine a match. Based on the result of this determination, the system verifies the identity of the person. This process is potentially vulnerable to "spoofing" attempts by fraudsters who try to fake their true identity by presenting an image of someone else's face to the biometrics identification system. The system needs to be able to determine whether the system captured an image of a real face or an image of a "spoof".

[0004] Current solutions for detecting such "impersonation"—that is, for making the assumption that an image is fake—rely on analyzing color images captured by a camera. However, this method has limited ability to overcome impersonation attempts that use video. Further complicating matters is the availability of image manipulation software that can be used to animate photographs. For example, there are mobile applications that can be downloaded to synthesize blinks. An imposter could present a mobile device displaying a photograph of their face to a biometric identification system while using such software to animate the photograph and make it appear as if someone else is interacting with the biometric identification system. This makes it more difficult for the facial biometric identification system to detect the imposter's attempt by requiring a live interaction with the person it is trying to identify.

[0005] During applications related to biometric enrollment processes or any other Know Your Customer (KYC) standards, determining whether a real human is present is particularly difficult when enrollment takes place in an unrestricted environment, for example, using a mobile phone at the user's location.

[0006] Where any prior art is referenced in this specification, please understand that such references do not constitute an admission that the prior art constitutes part of the common technical knowledge in the field in other countries. [Prior art documents] [Patent Documents]

[0007] [Patent Document 1] PCT / AU2024 / 050111 [Overview of the Initiative] [Means for solving the problem]

[0008] In a first embodiment, a liveness detection method for determining whether a face presented to a user is a real face is disclosed, the method comprising the step of acquiring an image sequence comprising a plurality of images. Each image comprises left-eye image data and right-eye image data. Each image in the image sequence is acquired from a face presented under its respective flash setting from a flash setting sequence, the flash setting of the flash setting sequence comprising at least two different flash setting values. The method comprises the step of processing the image sequence to detect speckles in the right-eye image data and left-eye image data from the image sequence, the speckles being caused by the eye reflecting a flash applied under a flash setting when acquiring at least one image of the sequence, and the flash setting for at least one image being brighter than the flash settings for the preceding and / or subsequent images. The method further comprises the step of determining whether a speckle in the left eye detected from the left-eye image data can be precisely matched with a speckle in the right eye detected from the right-eye image data, or vice versa. The method comprises the step of outputting a liveness detection result based at least in part on the determination.

[0009] In some embodiments, the processing step includes applying a filter to enhance spotted image data corresponding to spots in an image sequence compared to non-spotted image data in the image sequence, and / or to suppress non-spotted image data compared to spotted image data.

[0010] In some embodiments, spot detection involves applying a temporal algorithm to enhance the difference between at least one image and a preceding and / or subsequent image, wherein the output image of the temporal algorithm is processed to identify the portion of the image within at least one image that is a spot.

[0011] In some embodiments, the identified spots are required to have a minimum brightness level.

[0012] In some embodiments, filters for enhancing spotted image data compared to non-spotted image data, and / or suppressing non-spotted image data compared to spotted image data, include brightness filters applied to the output image or a portion of the output image from a temporal algorithm to apply a brightness threshold.

[0013] The brightness threshold may be selected such that only one spot exists for each eye.

[0014] In some embodiments, spot detection involves applying positional analysis, wherein the detected spots are required to be located within the user's cornea in the image sequence.

[0015] In some embodiments, one or more flashes applied according to a flash sequence have a predetermined shape and / or dimensions, or a configurable shape and / or dimensions.

[0016] One or more flashes applied according to a flash sequence may have one or more flash portions.

[0017] In some embodiments, spot detection involves applying a geometric analysis, wherein the detected spots are required to satisfy one or more geometric constraints, and the geometric constraints include constraints based on one or more of shape, aspect ratio, and contour.

[0018] In some embodiments, determining whether the left spot and the right spot coincide with each other includes determining whether the positions of the left spot and the right spot are mirror images of each other.

[0019] Determining whether the left spot and the right spot coincide with each other may further include determining whether the left spot and the right spot have a coincidental geometry.

[0020] In some embodiments, the method further includes the step of detecting spots for the left eye and the right eye from an image sequence, and calculating a measure of confidence associated with the determination if the detected spots for the left eye and the detected spots for the right eye are determined to be a match.

[0021] In some embodiments, the step of calculating a measure of confidence includes: iterating through the processing of an image sequence for spot detection one or more times, with each iteration the threshold applied by the filter being increased; stopping the iteration when it is no longer possible to detect one left eye spot in the left eye image data and one right eye spot in the right eye image data; and determining a measure of confidence based on the threshold applied during the last iteration of processing the image sequence for spot detection, when it was possible to detect one left eye spot in the left eye image data and one right eye spot in the right eye image data.

[0022] In another embodiment, the following apparatus is disclosed herein: an apparatus configured to determine whether a face presented with respect to a user is a real face, and comprising a processor configured to execute machine instructions for performing the method described above.

[0023] In another embodiment, a method for determining the identity of a subject by biometrics is disclosed herein, comprising the steps of: determining whether a presented face of the subject is a real face by the method described above; providing a two-dimensional image obtained from the presented face for biometric identification of the subject if the presented face is determined to be a real face; and outputting the result of the biometric identification.

[0024] In a further aspect, there is provided a computer-readable medium storing machine-readable instructions adapted to perform any of the methods described above when executed.

[0025] These and other aspects and embodiments of the present disclosure will become apparent to those skilled in the art based on the detailed description of various embodiments and / or implementations made with reference to the drawings, and then a brief description of the drawings is provided below.

[0026] Hereinafter, embodiments will be described merely by way of example with reference to the accompanying drawings.

Brief Description of the Drawings

[0027] [Figure 1] It is a high-level flowchart showing a liveness detection method according to an embodiment of the present invention. [Figure 2] It is a flowchart showing a process applied to a series of images to determine whether image data from the series of images shows the eyes of a real person according to an embodiment of the present invention. [Figure 3-1] It is a diagram conceptually showing a processing operation performed on an input image to identify possible spots according to an embodiment of the present invention. [Figure 3-2] It is an image showing the output from the temporal edge filter shown in FIG. 3-1. [Figure 3-3] It is an image showing the result of applying a flash filter to the image shown in FIG. 3-2. [Figure 4-1] It is a diagram showing a liveness detection process using a temporal flash filter according to an embodiment of the present invention. [Figure 4-2] It is a diagram showing a liveness detection process using a temporal flash filter according to another embodiment of the present invention. [Figure 5]This figure shows an exemplary program flow for implementing a liveness detection process according to an embodiment of the present invention, illustrating calls made to the flash processing module during the process detection and refinement phases. [Figure 6] This figure schematically illustrates an example of a system for authenticating or registering travelers according to an embodiment of the present invention. [Figure 7-1] The input image shows possible spots around the eyes and facial landmarks. [Figure 7-2] The input image shows possible spots around the eyes and facial landmarks. [Figure 7-3] This is the output image from a temporal processing applied to an image sequence including the image from Figure 7-1. [Figure 7-4] This is the output image from a temporal processing applied to an image sequence including the image from Figure 7-2. [Modes for carrying out the invention]

[0028] The following detailed description includes references to the accompanying drawings, which form part of the detailed description. The exemplary embodiments shown in the drawings and described in the detailed description are not intended to be limiting. Other embodiments may be used and other modifications may be made without departing from the spirit or scope of the subject matter presented. It will be readily apparent that the aspects of the disclosure described herein and shown in the drawings can be arranged, replaced, combined, divided, and designed in a wide variety of different configurations, all of which are envisioned in the disclosure.

[0029] Disclosed are methods and systems for detecting impersonation attempts or attacks. An impersonation attempt involves presenting a facial image or model, rather than a real face, to an automated system that uses facial biometrics for purposes such as enrollment, registration, or identity verification, in an attempt to deceive the automated system.

[0030] The "fakes" presented to automated systems in impersonation attempts may be static two-dimensional (2D) fakes, such as printouts or cutouts of photographs, or dynamic two-dimensional fakes, such as videos of faces displayed on a screen. Fakes may also be static three-dimensional (3D) fakes, such as static three-dimensional (3D) models or 3D renderings. Another type of fake is a dynamic 3D fake, such as a 3D model with face expression dynamics within a real or virtual camera.

[0031] Embodiments of the present invention may be used as standalone methods, i.e., methods not combined with other methods for detecting the presence of a genuine human being. Alternatively, embodiments of the present invention may be used in conjunction with other methods. This may help improve the overall robustness of the algorithm in a wider range of situations. For example, the applicant's application PCT / AU2024 / 050111, filed on 16 February 2024 and published on 22 August 2024 as WO2024168396, describes a method for detecting forgeries. The entirety of the aforementioned application is incorporated herein by reference.

[0032] Aspects of this disclosure are described herein using examples in the context of anti-spoofing for biometric identification of persons, such as passengers in air transport or other travelers. However, the disclosed technologies are applicable to any other automated systems that utilize facial biometrics.

[0033] In the context of air travel, passenger facial capture and biometric analysis can occur at various points during flight, such as check-in, baggage drop-off, security, and boarding. For example, in a typical identification system that utilizes facial biometric matching, the identification system includes image analysis algorithms that rely on color images captured by a camera. Therefore, systems using these algorithms have limited ability to detect when a pre-recorded or synthesized image sequence (i.e., a video sequence) rather than a real human face is presented to the biometric identification system's camera. The challenge becomes even greater when 3D fakes are presented.

[0034] Therefore, anti-spoofing measures for such systems may be implemented by estimating whether the image being analyzed is likely to have been obtained from a fake or from a real face, that is, by configuring those systems to estimate the liveness of the presented face, or by combining those systems with systems configured to do so.

[0035] Embodiments of the present invention provide a method for estimating the liveness of a face presented in a facial biometric system, that is, for determining whether it is a real face or a fake face. The disclosed method can be implemented as an anti-spoofing algorithm, step, or module in a facial biometric system. The system may be configured to enroll or register passengers, or to verify the identity of passengers, or both. The disclosure also encompasses a facial biometric system configured to carry out the method.

[0036] As described herein, embodiments of liveness detection methods utilize the reflective nature of the eye to detect whether an image presented to a system was taken from a real person. If a real person is present and interacting with the system, it is expected that an image taken from at least the person's eye when a “flash” (either a camera flash or a screen displaying a bright shape) is on will have image data in the image data indicating the presence of a flash in both the “eyes.” Since eyes are generally reflective, it is expected that turning on a flash will cause light to reflect into the eye, and the reflection is expected to be detectable as a spot in the image data. Thus, embodiments of liveness detection include glitter detection.

[0037] Figure 1 is a high-level flowchart illustrating a liveness detection method 100 according to an embodiment of the present disclosure. In step 102, the system acquires image data including multiple images representing both of the user's eyes with a predefined flash setting.

[0038] Multiple images may, but not necessarily, be captured in succession, for example, as a burst. For example, the user may be given time to rest or blink, or to receive information or instructions between the acquisition of separate images. Each image in a series of images is captured with a pre-set flash setting. The flash setting may be binarized and include on and off settings. The flash setting may have multiple levels. For example, when a selfie camera is used, a screen flash is used to provide flash. The brightness of the screen flash may be represented as 0 to 255, for example, in a system that uses 8-bit brightness encoding for pixels. The flash settings for multiple images may be considered a flash setting sequence.

[0039] In step 104, the acquired image data is processed to detect the user's eyes within the image data and to determine whether the image was obtained from real eyes. The output from the determination in step 104 may be used directly to output the final liveness determination result in step 106. Optionally, the output from the determination in step 104 may be combined with the results from one or more other liveness detection algorithms in step 108, and the combination is used to provide the final liveness determination in step 106. The box representing the combination step 108 is shown with a dashed line to indicate the optionality of this step. The image acquisition in step 102 and the processing in step 104 may be an interactive process in which the system outputs instructions to tell a passenger attempting to register or authenticate their identity to take a specific action.

[0040] A flash setting sequence may include a series of "on" (flash on) or "off" (flash off) settings. The minimum number of settings in a sequence is two. For a 3-setting sequence, the sequence may be "off-on-off" or "on-off-on". Capturing a user's image when the flash setting is controlled to alternate between "on" and "off" settings is expected to cause a visible flicker in the eye, detectable thanks to the reflective nature of the eye. Therefore, processing may be applied to determine whether the image data exhibits features expected in response to changes in the flash setting, namely, flash-reflecting light or flicker spots that are visible in only one or more of the acquired images but not (or equally not) visible in the others. In other implementations, there may be one or more flash settings between the extremes of the "on" and "off" settings, as long as the difference between the settings is still sufficient to expect the eye image data to have a distinguishable difference based on the applied setting. "Flash" does not necessarily refer to the camera flash. In embodiments where a selfie screen is used, a “selfie flash” is provided by increasing the brightness of the screen, i.e., by a screen flash. Thus, “flash” as referred to in these disclosures may be generalized to refer to an increase or significant change in brightness level that is output by the system when an image is taken and directed towards the user. This has the effect of illuminating the user’s face at a higher brightness level. Furthermore, a screen flash does not necessarily involve a sudden increase in the overall brightness of the screen. For example, an algorithm may present a particular bright shape.

[0041] Figure 2 illustrates one embodiment of the processing applied to a series of images to determine whether the image data from the series represents a real human eye. Each image in the series contains image data of an eye. Therefore, the images may be face images showing an eye, or alternatively, partial face images showing an eye. In step 202, image processing may be performed to detect the eye in order to identify the eye or a specific part of the eye (e.g., the cornea) within each image.

[0042] In step 204, the input sequence of images is spatially aligned with one another to establish a spatial correspondence between the images. This step is particularly useful in embodiments that do not employ deep learning techniques to extract or classify “spots” from the image data. In embodiments that utilize computer vision processing techniques, the alignment step 204 is performed to co-register the same features or landmarks in the images (e.g., iris, lens of the eye, etc.). When performing image alignment, the reference image is selected to minimize the expected eye movement between the reference image and the target image. For example, in an implementation where a sequence of three images is acquired for processing, the middle image may be selected as the reference, and therefore the first and third images are the target images. In embodiments that use neural network techniques, such as deep learning based on training data, step 204 may be omitted.

[0043] The reference image is expected to have been acquired with different flash settings than those used to acquire the target image. In step 206, the image is processed to enhance the effect of the flash on the image data. This may be done using temporal processing. Temporal processing may include calculating a difference image, but other types of temporal processing may be used, as can be determined by those skilled in the art. For example, higher-order temporal processing may be assumed. In one implementation, a temporal edge filter is applied to the image to strengthen the middle image compared to the other images and / or weaken the other images compared to the middle image. The filtering helps to strengthen the difference between the image acquired when the flash was on and the image acquired when the flash was off, and thus enhances the effect of the flash on the image data. The image obtained as a result of temporal processing (see, for example, Figure 3-1) is expected to show flash-induced spots. However, in practice, the output image from temporal processing may contain noise that is not necessarily flash-induced spots.

[0044] Preferably, in step 208, the output image from the temporal processing is processed by applying one or more further filters. This may help to remove noise from the output image or reduce the amount of noise in order to better separate the parts of the image in the output image from the temporal processing that show flash-induced spots. For example, the output image from the temporal processing (see, for example, Figure 3-1) is further processed to produce a binarized image (see, for example, Figure 3-2). This may be achieved in step 208 by applying a pixel brightness threshold to retain only pixels that have a brightness or intensity above the flash filter threshold. This may be considered as applying a flash filter at the flash filter threshold. In some embodiments, the output of the flash filter may be binarized, where all pixels with brightness values ​​above the flash filter threshold are assigned a first value, and all pixels below the flash filter threshold are assigned a second value.

[0045] Further processing may include extracting spots that satisfy constraints based on specific geometry, such as shape, size, and aspect ratio. This processing may be considered to provide a contour filter. Spotted image portions may be determined by identifying image portions that have a expected shape or contour. For example, when a user is taking an image using a selfie camera, spots resulting from the screen flash may be expected to have a specific shape corresponding to the shape of the screen flash, which will most often be rectangular. Therefore, shape processing or contour processing may be performed on candidate spots to further filter them and remove possible noise. It is assumed that, with sufficient camera resolution and processing speed, the shape of the screen flash and, therefore, the spots it produces, may be variable. Sufficiently high camera resolution may allow corneal details to be captured, in which case more complex shapes projected by the screen flash may be utilized. This allows for further ways to add the variability that impersonation attempts must satisfy, making it more difficult to deceive the system.

[0046] The spotted image portion may be extracted by identifying only the spotted image portion that exists symmetrically on both sides. That is, if a spotted image portion appears in one eye but a "matching" spotted image portion in the mirror image position of the other eye cannot be found, the spotted image portion is removed, i.e., filtered out. In this case, there are no matching spots in the left and right eyes.

[0047] The spots on the left and right eyes, considered to "match," may be required to have symmetrical positions. For example, symmetrical positions may mean the positions of the left and right eye spots relative to the respective inner corners of the eyes, mirror images of each other. In some embodiments, in order to be considered to "match," the spotted image portions may further need to have geometry that matches in one or more respects, such as shape and size, aspect ratio, etc.

[0048] To further isolate (i.e., extract) the spotted image portions, one or more of the processes described above may be included in step 208. If multiple processes are used, the order in which they are applied may be modified as configurable by a person skilled in the art. For example, a person skilled in the art may choose to apply the processing order that results in the least overall processing burden.

[0049] In step 210, a localization determination may be performed to determine the location of the spots. The location assigned to the candidate image portion may be the nearest mesh node in the mesh generated from the eye image data (or, depending on the implementation, the face mesh data). The location may be a specific biometric landmark. In embodiments in which an image alignment step is performed, the location may be defined using the coordinate system of the mutually aligned images.

[0050] In step 212, the identified spotted image portions are further filtered based on their location to determine whether they are likely to be spots resulting from a flash. This can help the entire algorithm determine which of the spots are most likely to appear as a result of a flash being on and reflected by a real person's eye ("flash detection"). For example, spots that are detected but located outside the reflective eye area may be a result of the user having oily skin or the user wearing glasses. By considering only spots located within the eye area that are likely to reflect a flash under the control of the liveness detection system, the system can better deal with noise in the data. In embodiments, the user may be instructed by the system to take a photograph or selfie with their face within a specific target area on the screen so that spots are expected to be seen within the corneal region of the eye. In this case, localization is further helpful in identifying spots located within the corneal region. Thus, the localization step 210 and the localization filtering step 212 also help to further filter out noise and identify spotted image portions. In some implementations, localization and location-based filtering may be performed earlier in the process, for example, before step 206 or step 208, or both.

[0051] Figures 7-1 to 7-4 show an example where spot extraction is noisy but improved by position-based processing. Figures 7-1 to 7-2 each show two input images from a series of images. Figures 7-3 to 7-4 show the results of temporal processing performed on the input images, which are further filtered by a flash filter. As can be seen, Figures 7-3 to 7-4 show many “bright” areas, but only one image portion within each eye is expected to represent a spot. By identifying the image portions within the corneal region of each eye, as represented by circles 701, 702, 703, and 704 in Figures 7-1 to 7-2, the number of image portions that may be extracted as “spots” is greatly reduced.

[0052] It will be understood that the order of processing may differ. For example, localization (step 210) and location-based filtering (step 212) may be performed before some or all of the denoising process in step 208. Steps 206 and 208 together may be considered a temporal flash process represented by the dashed box 214.

[0053] The process 200 described may be further modified. For example, an initial image may be captured with the flash set to "off" and processed to determine whether "spots" are present within the eye area, particularly within the reflective area of ​​the eye. This helps the system to verify the influence of ambient lighting, which may contribute to spots that should not be considered when evaluating liveness. The initial image may be an image taken as a separate capture before the series of images are taken. Alternatively, if acquired when the flash is off, the first image in the series may be used to determine whether there are any detectable spots caused by ambient light.

[0054] Figure 3-1 conceptually illustrates an example of the processing performed on an input image to identify possible spots. The processing is performed on images 302, 304, and 306, which are the first, second, and third images in a series of right-eye images, respectively, and the order of the images reflects the order in which they were acquired. Although not shown, the same processing is also performed on the left-eye images corresponding to each of the shown right-eye images.

[0055] In this example, the images were taken with the flash on, off, and on, respectively. The middle image, i.e., the second image 304, is used as the reference image, and the other two images 302 and 306 are taken as target images to be aligned with the reference image. The reference image 304 and the resulting aligned images 308 and 310 are provided as inputs to the temporal edge filter 312. Since different flash settings are used for the reference image 304, its image data is expected to be different from the data of the other two images. In this example, the temporal edge filter is configured to amplify the difference between the middle image 304 and the aligned images 308 and 310 by applying a larger weight (2x) to the middle image compared to the weight (1x) applied to the aligned images 308 and 310. It will be understood that the exact weighting applied is not a limiting factor. The output of the time-lag images using the three input images is shown in Figure 3-2.

[0056] To further filter the difference image 314, a threshold filter ("flash filter") 316 is applied to the resulting different image 314, so that only the portion of the image with brightness exceeding the threshold level of the flash filter 316 remains. This helps to separate possible spots, as shown in the resulting image 318 from the threshold filter shown in Figure 3-3. To determine the location of the possible spots, a localization algorithm 320 is applied to the outputs from the temporal filter 312 and the flash filter 316. There are several ways in which the location of the possible spots may be represented. The location may be determined based on the x,y coordinates of the possible spots. The determined location may be represented as a landmark associated with the eye, or as a mesh node in a mesh to represent the eye.

[0057] More broadly, Figure 3 illustrates the process by which input images are aligned relative to each other, passed through a temporal flash filter, and then spotted image portions located in specific regions of the eye, such as the corneal region, are identified from the output of the temporal flash filter.

[0058] Figure 4 shows a liveness detection process 400 utilizing a temporal flash filter according to an embodiment of the present disclosure. In this example, the temporal flash filter is applied iteratively until the algorithm terminates the iterative processing. The flash threshold at which termination occurs indicates a confidence level related to the determination that a spot caused by a forced flash has been detected and indicates the presence of a real human.

[0059] In step 402, an input image is acquired, similar to the acquisition of the image sequence in step 102 described above. Depending on the original image input, further processing to identify an eye or corneal region may be performed, for example, using segmentation or classification techniques. In step 404, the images are aligned, for example, by the alignment described above in relation to Figure 3-1. In embodiments utilizing artificial intelligence techniques, such as providing an object identifier or object classifier, the alignment step may not be necessary. In step 406, temporal filtering is applied to highlight the effect of forced flash. The output of the temporal filtering is passed to step 408 for further filtering to detect spots caused by flash. Step 408 may include one or more of the filters described above in relation to step 208. In step 410, a “count filter” is applied, which includes rejecting previously found spots if the number of spots found is not as expected. For example, in a process where only one spot is expected in each eye, the effect of including a count filter in process 400 is that the flash filter is applied one or more times, increasing the flash filter threshold while the processed images still have two or more spots, until each processed image has only one spot. This allows for the determination in step 412 whether there are matching spots in the processed images of the left and right eyes. If no matching spots are found, the system terminates process 400, and the algorithm determines that no spots attributable to flash were detected (step 414).

[0060] If it is determined that matching spots are found from the left and right eyes, the process then determines a confidence scale associated with this determination by applying the process in box 416. In step 418, the highest threshold applied so far in process 400 is raised. The amount of the raise may be set by the algorithm designer as needed. In step 420, a filter is applied again to the relevant image data, which in this case is the output of the temporal process, but with the raised threshold. In step 422, the algorithm determines whether a single matching spot exists from the left and right eyes. If it does, steps 418 to 422 are repeated. If it is determined in step 422 that a single matching spot does not exist from the left and right eyes, the system determines a confidence scale and terminates the process (step 424). The confidence scale is determined based on the highest threshold applied in which matching spots were found from the left and right eyes.

[0061] In some scenarios, it is expected that matching spots will only be found when the filter threshold used is sufficiently high to remove adequate noise. Therefore, the system may be configured to terminate process 400 only if no matching spots are found at any of the filter thresholds. For example, Figure 4-2 shows a variation of process 400. In this embodiment, the process includes steps 402 to 408 described above. After the application of further filters in step 408, the system determines whether the number of spots detected in each eye is the expected number (step 412 above).

[0062] From step 412, if the expected number of spots are found in each eye, the system determines that spots have been found and records the last applied flash filter threshold (step 430). The system then determines whether the last applied flash filter threshold has reached a predetermined maximum (step 432). If the expected number of spots are not found in each eye, the system proceeds directly to step 432. If the last applied flash filter threshold has reached the maximum value, the system terminates process 400 and, if any, determines a measure of confidence associated with detection based on the recorded threshold (step 434). If the last applied flash filter threshold is still below the predetermined maximum value, the system repeats steps 418 through 420 described above. This includes raising the flash filter threshold applied to the results of the temporal processing (step 418) and applying one or more further filters, including a flash filter at the raised threshold (step 420). The system then determines again whether the expected number of spots are found in each eye (step 412), and if so, updates the threshold at which spots are detected (step 430).

[0063] The exact algorithms for carrying out the above may vary. For example, a particular program may vary in how the temporal flash processing algorithm should be integrated into the processing loop as a module block, or how the termination conditions should be integrated into the processing loop. The algorithm may also vary (where applicable in the embodiment) based on the shape of the presented screen flash, or more broadly, the variety of the sequence of events presented to the user. In the usual sense, brightness thresholds or contour extraction may be applied to extract information from the image sequence that is expected to be extracted when a real face is presented (e.g., the timing or specific shape of the “flash” event).

[0064] An example is shown in Figure 5. In this example, a single matching spot is expected to be produced by the flash. It will be understood that the specific arrangement here is provided as an example and is not intended to limit the scope of this disclosure. The “start” 502 shown in Figure 5 may be thought to represent the start of a process relating to an input facial image of a user (e.g., a passenger). This may include the alignment of the input set of images and an initial pass of temporal flash processing on the aligned images. The aligned images are passed to a step 504 for identifying the location of the cornea in the image before iterative flash processing 505 is applied. This may help reduce the processing load by ignoring results that are not in the corneal region. Variations may exist. For example, the start 502 may include only the alignment of the images, and the step for identifying the location of the cornea is performed on the aligned images. In this case, the initial pass of temporal flash processing also includes the application of a temporal filter.

[0065] The iterative process 505 begins with the detection phase of the process. In the detection phase, in step 506, the system determines whether a single spot matching both eyes exists. The system may perform an initial pass of the flash process (block 512) to reach this determination. For the initial pass, the flash filter threshold is set to a predetermined minimum threshold. If a matching spot is found, the matching spot becomes the “first spot”. Box 508 represents the stage of process 500 where it is declared that the “first spot” has been found. If the first spot is not identified, one or more further calls to the temporal flash process 512 are made, with the flash filter threshold being raised with each call until the first spot is detected (pass 513). If the first spot is not detected and the flash filter threshold reaches a predetermined maximum value, the system determines that the “first spot” cannot be found (box 526) and the process terminates. Thus, the result is that no spots attributable to flash are found at any of the minimum applied flash thresholds, from the minimum to the maximum applied flash threshold, and therefore no liveness is detected.

[0066] If the first spot is identified (box 508), the system enters a “refinement” phase to further determine the confidence level associated with spot detection. In the refinement phase, in step 510, the system determines whether a spot was found after the most recent iteration of flash filtering. If positive, the system makes a call to the flash filtering module 512 using the raised flash filtering threshold (pass 515) and checks the result again to confirm that a spot was found with the raised flash filtering threshold (step 510). When the system enters the refinement phase from the detection phase for the first time, the output from the determination in step 510 is a positive determination. If, after a pass of flash filtering 512, matching spots are still detected in both eyes during the refinement phase (i.e., the result of the determination in step 510 is “yes”), the flash filtering is repeated by calling the flash filtering module 512 again (pass 515). If no more matching spots are found after a particular pass of flash filtering 512, the system terminates process 500 (pass 517).

[0067] When exiting the refinement phase of process 500, the system determines a measure of confidence based on the most recently used parameters of the flash filter used, which would have allowed for the extraction of matching spots. If an exit occurs from the detection phase, this means that no matching spots were extracted from processing block 505, and process 500 determines that no spots resulting from forced flash were detected.

[0068] Therefore, in the above, the subsequent flash processing 512, performed while processing 500 is in the refinement phase, is configured to determine the “strength” of the extracted spots. The “strength” may be defined by the brightness of the spots. However, it may instead be defined by the contrast between the spots and their neighboring pixels. Accordingly, the threshold is set to be applied, for example, to the absolute brightness of the output of the temporal edge filter, or to the contrast between the spots and surrounding pixels within the temporal edge output. In another example, the contrast may be the difference in brightness between image pixels at the same corneal position where the spots are found, captured at different time points and with different flash settings. This may be thought of as the “temporal brightness difference” associated with a particular spot.

[0069] In the above, the flash processing module 512 is configured to raise the flash filter threshold (step 514). Thus, each time flash processing 512 is performed, the flash filter threshold is raised from the most recently applied threshold. The flash processing module 512 further includes a spot filtering block 516 configured to apply one or more of the filters discussed above in relation to Figure 2, including the flash filter to be applied with the updated flash filter threshold. In this example, the spot filtering block 516 includes identifying the shape (or contour) of the portion of the image remaining from flash filtering and analyzing the shape (or contour). The spot filtering block 516 may also include analyzing the location of the spots and further determine whether the spots in the left eye and the spots in the right eye "match" each other.

[0070] In embodiments where the user is interacting with the screen, the intensity of the applied “flash” may depend on the brightness of the user’s screen, which may be set to a high brightness, a low brightness, or an automatically adjusted brightness level according to the user’s preference. The difference in brightness helps to extract information from the flash, even at low screen brightness settings.

[0071] Furthermore, ambient light settings can also affect flash intensity. Because users are required to frame their faces within limited yaw, pose, and roll, the camera (typically a selfie screen from a mobile device) needs to be held near the center of the face, resulting in flickering in both corneas. This flickering appears as spots in the eye image data. In the embodiments described herein, the measure of confidence is determined based on the maximum flash threshold used to extract two matching spots, one from each of the left and right eyes. Since the extraction process takes into account factors such as the geometry and location of the spots, the measure of confidence can be considered to provide a measure of the brightness, clarity, and shape of the flickering in both corneas.

[0072] Further robustness may be added to liveness detection systems that utilize the spot extraction-based processing described above. For example, randomization may be introduced into the capture pipeline to help prevent attackers from preparing recorded routines to deceive the capture process. Randomization may incorporate a face position randomizer that requires the user to place their face in a random position each time, making it more difficult for attackers to prepare face positioning routines. Randomization may also incorporate a flash display randomizer that randomizes the timing of flash sequences, making it more difficult for attackers to create face flicker routines. The flash display randomizer may further randomize the shape or contour of screen flashes if hardware is used that can acquire image data at a resolution high enough to process the data.

[0073] Further robustness may be added by incorporating one or more other liveness detection methods, such as the liveness detection method disclosed in PCT / AU2024 / 050111, the contents of which are incorporated herein by reference. For example, liveness detection including spot extraction disclosed herein is considered robust against two-dimensional static or dynamic impersonation, and three-dimensional static impersonation. Smile detection disclosed in the aforementioned application is considered robust against static two-dimensional or three-dimensional impersonation. Therefore, a combination of both liveness detection methods in a liveness detection system is useful against static two-dimensional and three-dimensional forgeries, as well as dynamic two-dimensional forgeries.

[0074] Figure 6 schematically shows an example of an automated system 600 intended for authenticating or registering travelers. The system 600 includes a device 602 which includes a camera 606 configured to acquire input image data and a flash output device 601, so that the device 602 can acquire images under a specific flash sequence as described herein. Alternatively, the device 602 is configured to receive a feed or images from an external camera configured to acquire images in a flash sequence. The device 602 includes a processor 603, which may be a central processing unit or other processor, configured to execute machine instructions to provide the liveness detection method according to this disclosure, either entirely or partially. For example, the processor 603 may be configured to execute only a portion of the method if a backend system with more powerful processing is required to process any of the steps of the method. The machine instructions may be stored in a memory device 607 located with the processor 603, as shown, or may reside partially or entirely in one or more remote memory locations accessible by the processor 603. The processor 603 may also access data storage 605 which contains the data to be processed and, in some cases, is adapted to store the results of the processing at least temporarily.

[0075] Device 602 further includes an interface arrangement 604 configured to provide audio and / or video interface capabilities for interacting with a user. Interface arrangement 604 includes a display screen and may further include other components such as a speaker and a microphone. A communication module 609 may also be present so that device 602 can receive or access data wirelessly, or transmit data or results to a remote location, such as a separate server computer, a monitoring station computer, or cloud storage, via a communication network that enables wireless communication 611.

[0076] During use, input image data is processed by a liveness detector 608 configured to perform a liveness detection method. As described above, the liveness detector 608 is provided as a computer program or module, which may be part of an application run by the processor 603 of device 602. Alternatively, the liveness estimator 608 may be supported by a remote server or be a cloud-based application accessible via a web-based application in a browser.

[0077] In Figure 6, the box representing device 602 is represented by dashed lines to conceptually illustrate that its components may be located within the same physical device or housing, or that one or more components may be located separately. For example, in embodiments where device 602 is a programmable personal device such as a mobile phone or tablet, the mobile phone or tablet may provide a single hardware device including an input / output (I / O) interface arrangement 604, a processor 603, data storage 605, a communication module 609, camera hardware 606, and local memory 607. Machine instructions for liveness detection can be stored locally, as previously mentioned, or accessed from the cloud or a remote location.

[0078] The application program for performing liveness detection may be provided as an application that runs on the local device's processor, for example, as a mobile application installed and run on a mobile phone. Alternatively, the application program may be provided to the local device as a web application via a regular web browser application (such as Chrome, Edge, or Safari) installed on the local device.

[0079] In particular, in the context of travel, the automated system 600 may be a kiosk terminal such as an airport check-in kiosk.

[0080] In some embodiments, device 602 is a “local device” because it is wirelessly connected to a backend system 612. Such a local device may be provided by a mobile phone or tablet. In the shown embodiments, the backend system 612 is a remote server or server system where the 1:N biometrics matching engine 614 resides. Communication between device 602 and the backend system 612 is represented by a dashed double arrow 611 and may be via a wireless network such as, but not limited to, a 3G, 4G, or 5G data network, or via a WiFi network. However, the backend system 612 may instead be provided by another server or server system, such as an airport server, which is separate from the server performing the 1:N matching but communicates with the server performing the 1:N matching.

[0081] In these embodiments, the backend system 612 may include a backend liveness detector 616 configured to perform either partially or completely the same method as that performed by the liveness detector 608. In this case, the camera feed data is also sent to the backend liveness detector 616. That is, while the device's liveness detector 608 is processing the live camera feed, the camera data is also being supplied to the backend server 612 for the same processing. This serves the purpose of performing a verification run of the processing to ensure that the results returned by liveness detection are not corrupted, or to perform steps of the liveness detection method that may be too computationally intensive for the local device 602 to handle, or both. The automated process for authenticating or enrolling travelers proceeds to perform 1:N matching only if the results from both the local liveness detector 608 and the backend liveness detector 616 both indicate "liveness" of the face images in the camera feed.

[0082] To reduce the likelihood of a video injection attack, the overall robustness of the system may be enhanced by applying video fingerprinting or encryption. A video injection attack can be carried out by a virtual camera playing a pre-recorded video that attempts to deceive the liveness check. Over time, the same video may be injected repeatedly, eventually succeeding by chance, which may well match the randomness of the capture process. To prevent an attacker from attempting to play the same video more than once, video fingerprinting may be applied to the image sequence captured during liveness testing. For example, when the input sequence is provided to local device 602, a fingerprint may be created on the sequence, or output data from a liveness detection module. If the same video is played later, the application will recognize the fingerprint, realize that the video may be injected in an injection attack, and prevent the transaction to the backend 612 from proceeding. Video fingerprinting may be carried out, for example, by creating a rolling hash function associated with the real-time video stored locally on device 602.

[0083] An attacker could potentially bypass the front-end capture system, for example, the mobile device 602, and inject a sequence of data into the backend 612 to successfully authenticate. Therefore, cryptographically protecting biometric data in transmission is particularly useful against injection attacks. In some embodiments, as an alternative to or in addition to video fingerprinting, the transmission of biometric data between the mobile device 602 and the backend 612 may be protected by cryptographic techniques. The backend 612 may be a security server. For example, data from a liveness module provided on the local device 602 to perform liveness detection may be encrypted using a locally stored public key, but require a private key on the backend server for decryption. The encrypted data is sent to the backend server 612, where it is validated. A message indicating whether the validation was successful is returned to the local device 602. This validation may be required for the liveness detection process to proceed.

[0084] In the context of travel, liveness detection is described as part of a check performed before biometric identification is carried out. However, the implementation of biometric identification does not affect the function of liveness detection and is therefore not considered part of the invention in any of the disclosed embodiments. For example, liveness detection may be performed in systems that do not perform biometric identification. For example, liveness detection may be performed in a system to check whether a person passing through or present at a checkpoint is using a “fake” to conceal their identity or impersonate someone else, for example, to join a video conference or register themselves in a particular user database.

[0085] The foregoing parts may be changed and modified without exceeding the spirit or scope of this disclosure.

[0086] For example, other processes may be utilized that provide the ability to detect the relevant eye area and the presence of spots resulting from the applied flash. As suggested above, a trained model may be used to determine whether or not spots resulting from the applied flash are found. A mixture of methods may be employed. For example, eye or corneal identification may be performed using a classifier, and the presence of spots resulting from the applied flash may be determined using temporal processing on bounding boxes identified by the classifier.

[0087] As another example, instead of presenting a single flash shape each time the flash is turned on, the flash may include a flash pattern having one or more flash portions, and therefore, when the flash is turned on, there may be a scenario in which one or more flash portions are “fired.” The shape of each flash portion may also be adjustable. Thus, in embodiments of this disclosure, the “flash” may consist of one or more flash portions. Thus, a single spot detected in the left and right eyes, forced by such a flash, consists of one or more spot portions corresponding to one or more flash portions of the flash. This can be useful in adding further randomization to the algorithm, particularly in implementations where the required resolution can be provided by image capture hardware.

[0088] This specification describes various embodiments with reference to numerous specific details, which may differ from implementation to implementation. Limitations, elements, characteristics, features, advantages, or attributes not expressly stated in the claims should not be considered essential or indispensable features. Accordingly, this specification and the drawings should be interpreted as illustrative rather than restrictive. In subsequent claims and in the preceding description, unless the context should be interpreted otherwise by express wording or necessary suggestion, the word “comprise,” or variations such as “comprises” or “comprising,” is used in an inclusive sense, that is, to specify the presence of a stated feature in various embodiments of this disclosure, but without precluding the presence or addition of further features. [Explanation of Symbols]

[0089] 100 Liveness Detection Methods 200 processes Images 302, 304, 306 Images 308 and 310, aligned to one another. 312 Temporal Edge Filter 314 difference images 316 Threshold filters, flash filters 318 Result Images 320 Location Algorithms 400 Liveness Detection Process 505 Iterative flashing 512 Temporal flash processing, flash processing module 516 Spot Filtering Block 600 automated systems 601 Flash Output Device 602 devices 603 Processor 604 Interface Configuration 605 Data Storage 606 Camera 607 Memory Devices 608 Liveness detector, liveness estimator 609 Communication Module 611 Wireless communication 612 Backend systems, backend servers 614 1:N Biometrics Matching Engine 616 Backend Liveness Detector 701, 702, 703, 704 yen

Claims

1. A liveness detection method for determining whether a face presented to a user is a real face, A step of acquiring an image sequence comprising multiple images, wherein each image comprises left eye image data and right eye image data, and each image in the image sequence is acquired from the presented face under the respective flash settings from a flash setting sequence, and the flash settings of the flash setting sequence comprise at least two different flash setting values. A step of processing the image sequence to detect spots in the right-eye image data and the left-eye image data from the image sequence, wherein the spots are caused by the eye reflecting a flash applied under the flash setting when acquiring at least one image of the sequence, and the flash setting for the at least one image is brighter than the flash setting for the preceding and / or subsequent images, The steps include determining whether the spot on the left eye detected from the left eye image data can be precisely matched with the spot on the right eye detected from the right eye image data, or vice versa, The steps include outputting liveness detection results based at least partially on the determination, and Methods that include...

2. The method according to claim 1, wherein the processing step includes applying a filter to enhance spotted image data corresponding to the spots in the image sequence compared to non-spotted image data in the image sequence, and / or to suppress the non-spotted image data compared to the spotted image data.

3. The method according to claim 1 or 2, wherein detecting the spots involves applying a temporal algorithm to enhance the difference between the at least one image and the preceding and / or subsequent images, wherein the output image of the temporal algorithm is processed to identify the portion of the image within the at least one image that is a spot.

4. The method according to claim 3, wherein the identified spots are required to have a minimum brightness level.

5. The method according to claim 3 or 4, wherein a filter for enhancing spotted image data compared to non-spotted image data, and / or suppressing said non-spotted image data compared to said spotted image data, includes a brightness filter applied to the output image from the temporal algorithm or a portion of the output image to apply a brightness threshold.

6. The method according to claim 5, wherein the brightness threshold is selected such that there is only one spot for each eye.

7. The method according to any one of claims 1 to 6, wherein the detection of the spot is required to apply positional analysis, and the detected spot is required to be located within the user's cornea in the image sequence.

8. The method according to any one of claims 1 to 7, wherein one or more flashes applied according to a flash sequence have a predetermined shape and / or dimensions, or a configurable shape and / or dimensions.

9. The method according to claim 8, wherein the one flash or the plurality of flashes applied according to the flash sequence have one or more flash portions.

10. The method according to any one of claims 1 to 9, wherein the detection of the spots is required to apply a geometric analysis, the detected spots are required to satisfy one or more geometric constraints, the geometric constraints include constraints based on one or more of shape, aspect ratio, and contour.

11. The method according to any one of claims 1 to 10, wherein determining whether the left spot and the right spot coincide with each other includes determining whether the positions of the left spot and the right spot are mirror images of each other.

12. The method according to claim 11, further comprising determining whether the left spot and the right spot coincide with each other, and determining whether the left spot and the right spot have a coincident geometry.

13. The method according to any one of claims 1 to 12, further comprising the step of detecting the left eye spot and the right eye spot from the image sequence, and if it is determined that the detected left eye spot and the detected right eye spot match each other, calculating a measure of confidence related to the determination.

14. The step of calculating the aforementioned confidence scale is A step of repeating the processing step of the image sequence for detecting the spots once or more times, wherein in each iteration the threshold applied by the filter is increased; The steps include: stopping the iteration when it is no longer possible to detect one spot on the left eye in the left eye image data and one spot on the right eye in the right eye image data; A step of determining the confidence measure based on the threshold applied during the last iteration of the processing step of the image sequence for detecting the spots, which made it possible to detect one spot on the left eye in the left eye image data and one spot on the right eye in the right eye image data. The method according to claim 13, including the method described in claim 13.

15. An apparatus configured to determine whether a face presented with respect to a user is a real face, comprising a processor configured to execute machine instructions to carry out the method according to any one of claims 1 to 14.

16. A method for determining the identity of a subject using biometrics, A step of determining whether the presented face of the subject is a real face by the method described in any one of claims 1 to 14, If the presented face is determined to be a real face, the step of providing a two-dimensional image obtained from the presented face for biometric identification of the subject, The steps include outputting the results of the biometric identification and Methods that include...

Citation Information

Patent Citations

  • Automated facial detection with Anti-spoofing

    WO2024168396A1