Automatic face detection with anti-spoofing

The method analyzes facial movements, expressions, and illumination responses to determine liveness, effectively addressing spoofing attacks in facial recognition systems, enhancing security by distinguishing between real faces and spoofing attempts.

JP2026506140APending Publication Date: 2026-02-20ICM AIRPORT TECHNICS AUSTRALIA PTY LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025547809
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-02-16
Filing Date
2024-02-16
Publication Date
2026-02-20

AI Technical Summary

Technical Problem

Existing facial recognition systems are vulnerable to spoofing attacks, particularly in environments where images or videos are presented to deceive identity, and existing methods fail to effectively detect the fraudster's attempts by requiring detection the fraudster's attempts by requiring detection the fraudster's attempts by presenting an animated photograph, making it appear as if that person is interacting with the biometric identification system, the system needs to be authenticated by presenting an image of someone else's face to the system. This makes it more difficult for facial biometric identification systems to detect the fraudster's attempts by requiring live interaction with the biometric identification process is addressed by the system. The challenge of existing technologies have not been effectively solved by requiring new technologies have not been effectively solved by existing technologies have not been effectively solved by existing technologies have not been effectively solved by existing technologies have not been effectively solved by existing technologies have not been effectively solved by existing technologies have not been effectively solved by existing technologies have not been able to detect spoofing attempts using video or 3D models.

Method used

The method involves analyzing facial movements, expressions, and illumination responses to determine the liveness of a face by instructing users to align their facial images with a randomly selected target, analyzing changes in facial feature metrics, and comparing contrast levels under different illumination levels, using machine learning and reinforcement learning models to estimate whether the presented face is genuine.

Benefits of technology

This method effectively distinguishes between real faces and spoofing attempts, enhancing the security of facial biometric systems by reducing the likelihood of successful impersonation, particularly in environments where static or dynamic images and 3D models are used.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026506140000001_ABST
    Figure 2026506140000001_ABST
Patent Text Reader

Abstract

A method is disclosed for estimating whether a presented face of a user is a genuine face by analyzing collected image frames of the presented face, the method including one or more of the steps of determining and analyzing movements made by the user to align the presented face image with a randomly selected target, determining and analyzing changes in facial feature metrics in the face image in response to changes in the user's facial expression, and determining and analyzing the effect of illumination changes on contrast levels between a region of the face and one or more regions adjacent to the region of the face in the collected images.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to automated facial recognition or identification systems, and in particular to systems that address the potential vulnerability of these systems to spoofing attacks in which someone attempts to be authenticated by presenting an image of their face to the system. [Background technology]

[0002] The use of biometrics in authenticating or identifying individuals has gained momentum in recent years, particularly given advances in facial recognition and image processing techniques. An application that could readily adopt such use is passenger identification or registration, particularly at airports, where self-service kiosks already exist that allow passengers to check in for flights, print boarding passes, or print baggage tags, among other functions. With advances in computing and camera technology, facial biometrics verification may also increasingly be used in other scenarios, such as building access control.

[0003] In a facial biometric identification system, an image of a person's face is captured, analyzed, and compared to a database of registered faces to determine whether there is a match. Based on the outcome of this determination, the system confirms the person's identity. This process is potentially vulnerable to "spoofing" the biometric identification system, i.e., an attempt by a fraudster to "impersonate" another person by presenting an image of someone else's face to conceal their true identity. The system needs to be able to determine whether it has captured a live face image or a "spoof" image.

[0004] Current solutions for detecting such "spoofing," i.e., inferring that an image is a spoof, rely on analyzing color images captured by a camera. However, this approach is limited in its ability to thwart spoofing attempts using video. Compounding the problem is the availability of image manipulation software that can be used to animate photographs. For example, there are mobile applications that can be downloaded to synthesize eye blinks. A fraudster may present a mobile device display showing a photograph of another person's face to a biometric identification system and simultaneously use such software to animate the photograph, making it appear as if that person is interacting with the biometric identification system. This makes it more difficult for facial biometric identification systems to detect the fraudster's attempts by requiring live interaction with the person they are trying to identify.

[0005] Where prior art is referred to herein, it is to be understood that such reference is not an admission that the prior art forms part of the common general knowledge in the art in other countries. Summary of the Invention [Means for solving the problem]

[0006] In a first aspect, the present invention provides a method for estimating whether a presented face of a user is a genuine face by analyzing collected image frames of the presented face, the method comprising: (a) determining and analyzing movements made by a user to align a facial image of a presented face with a randomly selected target; (b) determining and analyzing changes in facial feature metrics in the facial image in response to changes in the user's facial expression; and (c) determining and analyzing the effect of illumination variations on the contrast level between the facial region and one or more regions adjacent to the facial region in the collected images; Contains one or more of:

[0007] In some examples, (a), (b), and (c) are performed sequentially.

[0008] In some examples, at least one of (a), (b), and (c) is performed simultaneously with another one of (a), (b), and (c).

[0009] In some examples, determining and analyzing movements made by the user to align the facial image of the presented face with the randomly selected target includes: displaying a randomly selected target on a display screen and instructing a user to perform a movement on the display screen to move the user's facial image from its current position so that the user's facial image matches the randomly selected target; analyzing image frames collected while the user is instructed to track a randomly selected target; estimating whether the presented face is a genuine face based on the analysis; and Includes.

[0010] The targets may be randomly chosen in that they have randomly chosen locations, sizes, or both.

[0011] In some examples, analyzing the image frames collected while the user is instructed to track the randomly selected target includes estimating movements made by the user.

[0012] In some examples, estimating whether the presented face is a genuine face includes comparing the determined one or more movements with movements that a person would be expected to make.

[0013] In some examples, the estimation is performed by a machine learning model.

[0014] In some examples, the estimation is based on reinforcement learning, trained using data on natural movements.

[0015] In some examples, determining the movement includes determining a path between the current location and a randomly selected target.

[0016] In some examples, determining and analyzing changes in facial feature metrics in the facial image in response to changes in the user's facial expression includes: instructing a user to make facial expressions that cause distortion of visually detectable facial features; analyzing image frames collected while the user is instructed to make the facial expression; estimating whether the presented face is a genuine face based on the analysis; and Includes.

[0017] In some examples, analyzing the image frames collected while the user is instructed to make the facial expression includes obtaining a time series of metric values ​​from the series of analyzed frames by calculating a metric based on the position of one or more facial features from each analyzed frame.

[0018] In some examples, estimating whether the presented face is a genuine face based on the analysis includes determining momentum in the time series of metric values ​​and comparing the momentum to an expected momentum profile for a genuine smiling face.

[0019] In some examples, the user is instructed to smile.

[0020] In some examples, for each analyzed frame, the calculated metric is or includes the ratio of the distance between the eyes of the detected face in the analyzed frame to the width of the mouth of the detected face.

[0021] In some examples, the method includes comparing image frames collected while the user is instructed to make a facial expression with a reference image in which the user has a neutral expression, and selecting from the analyzed image frames an image that is most similar to the reference image as an anchor image in which the user is considered to have a neutral expression.

[0022] In some examples, analyzing the effect of variations in illumination on presented faces in the image data includes: capturing one or more first image frames of the presented face at a first illumination level and capturing one or more second image frames of the presented face at a second illumination level different from the first illumination level; analyzing the first image frame to determine a first contrast level between a detected face region in the first image frame and an adjacent region adjacent to the detected face region; analyzing the second image frame to determine a second contrast level between the detected face region in the second image frame and an adjacent region adjacent to the detected face region of the second image frame, wherein the relationship between the adjacent region of the second image frame and the detected face region is the same as the relationship between the adjacent region of the first image frame and the detected face region; comparing the first and second contrast levels to estimate whether the presented face is likely to be a genuine face; Includes.

[0023] In some examples, comparing the first and second contrast levels includes determining whether a change between the first contrast level and the second contrast level is greater than a threshold value.

[0024] In some examples, during the step where the user is requested to align the user's facial image with a target area on the screen, the method further comprises: prompting the user to align a facial image of the user's presented face with the first, larger target area; Upon detecting a face image within the first target area, setting an area generally encompassed by the face image as a reference target area. Includes.

[0025] In some examples, the method includes applying a tolerance range near the first, larger target area, whereby detecting a facial image within the tolerance range triggers setting the area of ​​the facial image as a reference target area.

[0026] In some examples, the method includes applying a tolerance around the reference target area such that a facial image of the user's presented face is considered to remain within the reference target area if it is within the tolerance around the reference area.

[0027] In another aspect, the present invention provides an apparatus for estimating whether a presented face of a user is a genuine face by analyzing collected image frames of the presented face, the apparatus comprising a processor configured to execute machine instructions implementing the above-described method.

[0028] In some examples, the device is a local device used or accessed by a user. In other examples, the device is a kiosk at an airport. The kiosk may be a check-in kiosk, a bag drop kiosk, a security kiosk, or other kiosk.

[0029] In some examples, the local device is a mobile phone or a tablet.

[0030] In some examples, the collected image frames are sent over a communications network to a back-end system and processed by a processor in the back-end system configured to execute machine instructions that at least partially implement the above-described method.

[0031] In some examples, if the results of processing by the device's processor and the results of processing by the backend system's processor both estimate that the presented face is a genuine face, the presented face is estimated to be a genuine face.

[0032] In some examples, the device allows the user to interface with an automated system, a biometric matching system, that enrolls or verifies the user's identity.

[0033] In some examples, the user is a passenger on an airline trip and the backend system is a server system that hosts a biometric matching service or is a server system connected to another server system that hosts a biometric matching service.

[0034] In another aspect, the present invention provides a method for biometrically determining a subject's identity, the method comprising: estimating whether the presented face of the subject is a genuine face according to the method described above; If the presented face is presumed to be a genuine face, providing the collected two-dimensional image of the presented face for biometric identification of the subject; outputting the result of the biometric identification; Includes.

[0035] In some examples, this method occurs during the check-in process by an air travel passenger.

[0036] In a further aspect, there is provided a computer readable medium having stored thereon machine readable instructions adapted, when executed, to perform any of the methods set forth above.

[0037] In a fourth aspect, there is provided a biometric identification system including an image capture device, a depth data capture device, and a processor configured to execute machine-readable instructions that, when executed, are adapted to perform the method for biometrically determining a subject's identity described above.

[0038] Embodiments will now be described, by way of example only, with reference to the accompanying drawings, in which: [Brief explanation of the drawings]

[0039] [Figure 1] 1 is a schematic diagram of a liveness estimation method according to an embodiment of the present invention; [Figure 2] 1 is a flowchart of a facial expression detection process according to one embodiment. [Figure 3(1)] It is an image of a person with a smiling expression. [Figure 3(2)] This is an image of a person with a blank expression. [Figure 4(1)] FIG. 10 shows the ratio of lip spread to eye separation (LD / ED) through the frame as the facial expression changes from neutral to smiling. [Figure 4(2)] FIG. 10 shows LD / ED through frames when the facial expression changes from smiling to neutral. [Figure 5(1)] FIG. 10 shows LD / ED and momentum in LD / ED calculated from image frames collected over a period of time while maintaining a neutral expression. [Figure 5(2)] FIG. 10 illustrates LD / ED and momentum in LD / ED calculated from image frames collected over a period of time when the expression is a slow smile made by a real face. [Figure 5(3)] 10A and 10B show LD / ED and momentum in LD / ED calculated from image frames collected over a period of time when the facial expression is a medium-paced smile produced by a genuine face. [Figure 5(4)]FIG. 10 illustrates LD / ED and momentum in LD / ED calculated from image frames collected over a period of time when the expression is a fast-paced smile made by a real face. [Figure 6] (1) A diagram showing a first image of a face with a neutral expression and a second image of a face with a smiling expression, with the face changing from a neutral expression to a smiling expression; (2) A schematic diagram showing the time at which image frames are expected to be collected when the camera frame rate is 10 frames per second (fps); (3) A diagram showing a time series of LD / ED data measured from 12 frames taken at 2 fps; and (4) A schematic diagram showing how the time series of Figure 6(3) is used to populate or interpolate values ​​to populate a data series for analysis at a target sampling rate higher than 2 fps. [Figure 7] FIG. 10 is a diagram illustrating a schematic user interface on a mobile device where a user is asked to match an image of the user's face to a randomly selected target. [Figure 8] FIG. 1 is a diagram of a reinforcement learning model. [Figure 9] FIG. 1 illustrates an exemplary facial motion test, according to one embodiment of the present invention. [Figure 10] FIG. 2 illustrates an exemplary facial dimension analysis according to one embodiment of the present invention. [Figure 11] FIG. 10 is a diagram showing an example of a face region and four adjacent non-face regions. [Figure 12] FIG. 1 illustrates an exemplary facial target fitting process according to one embodiment of the present invention. [Figure 13(1)] FIG. 13 shows a schematic diagram of the rough fitting step mentioned in FIG. 12. [Figure 13(2)] FIG. 13 is a diagram illustrating the simplified fitting step mentioned in FIG. 12. [Figure 13(3)] FIG. 13 is a diagram illustrating the reference reset step mentioned in FIG. 12; [Figure 14]FIG. 1 is a diagram illustrating a schematic example of an automated system for authenticating or registering travelers, according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0040] In the following detailed description, reference is made to the accompanying drawings, which form a part of the detailed description. The exemplary embodiments described in the detailed description and illustrated in the drawings are not intended to be limiting. Other embodiments may be utilized, and other changes may be made, without departing from the spirit or scope of the presented subject matter. It will be readily understood that aspects of the present disclosure, as generally described herein and illustrated in the drawings, can be arranged, substituted, combined, separated, and designed in a wide variety of configurations, all of which are contemplated in this disclosure.

[0041] Disclosed are methods and systems for detecting spoofing attempts or attacks in which an image or model of a face, rather than a live face, is presented to an automated system that uses facial biometrics for purposes such as identity enrolment, registration, or verification in an attempt to fool the automated system.

[0042] The "spoofs" presented to an automated system in an impersonation attempt may be static two-dimensional (2D) spoofs, such as a photographic printout or cutout, or dynamic 2D spoofs, such as a video of a face presented on a screen. The spoof may be a static three-dimensional (3D) spoof, such as a 3D model or a 3D rendering. Another type of spoof is a dynamic 3D spoof, for example, a 3D model with facial expression dynamics in a real or virtual camera.

[0043] Aspects of the present disclosure are described herein in the example context of anti-spoofing for biometric identification of persons, such as airline transit passengers or other travelers, although the disclosed techniques are applicable to other automated systems that utilize facial biometrics.

[0044] In the context of air travel, capture and biometric analysis of passengers' faces may occur at various points, such as flight check-in, baggage drop, security, and boarding. For example, in common identification systems that utilize facial biometric matching, the identification system includes image analysis algorithms that rely on color images captured by a camera. Systems that use these algorithms are therefore limited in their detection capabilities when pre-recorded or synthesized image sequences (i.e., video sequences) are presented to the biometric identification system's camera, rather than a person's authentic face. The challenge becomes even greater when 3D spoofs are presented.

[0045] Therefore, spoofing prevention for such systems can be achieved by configuring them, or combining them with systems configured to estimate whether the image being analyzed is likely to represent a spoof or a real face, i.e., estimate the liveness of the presented face.

[0046] Embodiments of the present invention provide a method for estimating the liveness of a face presented to a facial biometric system, i.e., determining whether it is a real face or a spoof. The disclosed method can be implemented as an anti-spoofing algorithm, step, or module in a facial biometric system. The system can be configured to enroll or register passengers, or to verify passenger identities, or both. The present disclosure also encompasses a facial biometric system configured to implement the method.

[0047] FIG. 1 is a high-level schematic diagram of a liveness estimation method 100 according to one embodiment of the present invention. Image data is received or collected by a system implementing the method in step 102. The image data is processed in step 104, and then liveness estimation is performed in step 106. In the illustrated embodiment, the processing of the image data includes facial expression analysis 108, motion tracking analysis 110, and facial lighting response analysis 112. The processing in step 104 may be an interactive process, where the system will output instructions to instruct a passenger attempting to register or authenticate their identity (e.g., at check-in) to take a specific action. Image data captured while the passenger is performing the action may then be analyzed. Arrow 105 represents the interactive process, where additional image data is collected or received and analyzed during processing (step 104).

[0048] It should be noted that in other embodiments, the processing performed in step 104 may differ from that shown in FIG. 1 by including only one or two of the three types of analysis shown in FIG. 1.

[0049] Facial expression analysis Facial expression detection 108 analyzes facial images detected in the input image data. The input image data is collected while a user is instructed to make a particular facial expression or series of expressions to determine whether facial expressions likely to be made by a genuine face can be detected. The facial expressions are of a type such that at least a partial set of the user's facial features or muscles are expected to move while the user is making the expression. The one or more movements cause a "distortion" in the facial features compared to a neutral or expressionless face. Thus, the analysis for performing facial expression detection 108 characterizes this distortion from the image data to determine whether a genuine face has been captured making the expression or whether the facial image is likely to be captured from a "spoof."

[0050] 2 is a high-level diagram of facial expression detection 108 according to one embodiment. In this embodiment, facial expression detection 108 is conceptually shown to occur after detecting facial images in the image data (step 113) (as represented by the dotted rectangular box). However, in other embodiments, facial image detection 113 may be included as part of the facial expression detection process 108.

[0051] 2, facial expression detection 108 in step 114 detects one or more facial features in each processed image frame. In step 116, one or more metrics may be calculated from the detected facial features, such as the width or height of a particular facial feature or the distance between facial features. In step 118, the positions of the detected features, or the metrics calculated from step 116 (if step 116 is performed), are tracked over a period of time during which the user is required to make that facial expression. In step 120, the time series of data is analyzed to determine whether it represents an authentic, "live" face or a spoof that has been incorporated into the image data.

[0052] An example of the process will now be described with reference to Figures 3(1) and 3(2). For illustrative purposes, Figures 3(1) and 3(2) show an example in which the user is instructed to produce a smile. The facial features identified are the person's eyes and mouth. When a person changes from a neutral or non-smiling expression (Figure 3(2)) to a smiling expression (Figure 3(1)), the distance between the eyes is expected to remain the same, but the mouth is expected to widen. Therefore, the ratio of the mouth width to the distance between the eyes is expected to increase, as shown in Figure 4(1). Conversely, when a person changes from a smiling expression to a neutral or non-smiling expression, this ratio is expected to decrease, as shown in Figure 4(2). Therefore, metrics calculated from the facial features may include the distance between the eyes (ED) and the distance across the lip width (LD), and the ratio of the two distances (i.e., LD / ED). In this example, the value of the ratio LD / ED is tracked over time, and the time series of values ​​of the metric is analyzed. It should be noted that in other implementations where the user is required to smile or make other facial expressions, different metrics may be tracked within the same premise of tracking movement or distortion in facial metrics across a series of images.

[0053] The above process may be generalized to include other embodiments. The generalization may be in one or more different ways. For example, facial expressions other than smiling may be used, as long as the expression is expected to produce a measurable change or "contortion" of facial muscles. The calculations used to quantify the change may include, for example, different numbers of facial points for the purpose of analyzing different types of facial expressions. The calculations may also be nonlinear, for example, including models based on nonlinear calculations such as polymodels or spline models rather than linear models.

[0054] In facial expression analysis 108, the analysis performed on the time series of values ​​(step 120) may include determining the momentum of the time series of values ​​and examining the momentum characteristics to see if those values ​​indicate that a "live" expression is being produced, i.e., a real face. Different methods may be used to measure momentum. One example is Moving Average Convergence / Divergence (MACD) analysis, although other tools for measuring momentum may be used instead.

[0055] Referring again to the example where the user is asked to smile and the metric measured over time is the LD / ED ratio, the momentum of the time series will differ depending on whether there is no smile (Figure 5(1)), or whether the smile is a slow smile (Figure 5(2)), a medium smile (Figure 5(3)), or a fast smile (Figure 5(4)). In Figures 5(1) through 5(4), the LD / ED metric (in %) is shown in the top graph. The MACD analysis is shown in the bottom graph of Figures 5(1) through 5(4). The horizontal axis represents the sample points.

[0056] As can be seen from Figure 5(1), when there is no change in facial expression, as might be expected from a static 2D or 3D spoof, the changes in the LD / ED ratio do not show consistent momentum or momentum trends. For example, in Figure 5(1), the amplitude of MACD line 501 is within a low threshold, in this case plus or minus 0.1. On the other hand, as can be seen from Figures 5(2) through 5(4), when a smile is produced by a genuine face (from a neutral expression), regardless of the pace of the smile, momentum generally increases toward the end of the time period, and MACD line 502 has a much larger amplitude. Determining whether one or more of the aforementioned characteristics can be observed from the momentum data helps indicate whether the smile is "animated," i.e., produced by a genuine face.

[0057] In the above, the momentum being analyzed indicates the momentum of a facial metric when a facial expression is expected to change to or from "neutral." Therefore, a facial expression analysis algorithm needs to have access to an image that can be considered to provide a neutral or expressionless face. This image is sometimes referred to as an "anchor" image. This may be done by asking a person interacting with the automated system to make a neutral face. The algorithm may set a threshold or threshold range for one or more metrics being analyzed and assign images that meet the threshold or threshold range as anchor images.

[0058] Alternatively, the anchor image may be selected based on a reference image of a person interacting with the automated system, in which the face is expected to have a neutral expression. The reference image may be a previously existing image, such as an identification photo; an example is a driver's license photo or a passport photo. For example, each image in a series of input images is compared to the reference image. The input image deemed "closest" to the reference image will be selected as the "anchor" image. Image frames in the series of input images after the anchor image may then be analyzed to determine whether the "face" in the image is a real face or a spoof. The comparison of the reference image and the input image may be performed using a biometric method. The comparison may be performed based on specific facial feature metrics used in facial expression analysis by comparing the metric(s) calculated from the reference image with the same metric(s) calculated from the input image, and identifying the input image from which the calculated metric(s) are closest to the metric(s) calculated from the previously existing image. The identified input image is then selected as the "anchor" image.

[0059] In an air travel scenario where a passenger is interacting with an automated system, for example to check in, the passenger's passport photograph may be used to provide a reference image. This has the advantage that, as explained above, passport photographs are expected to generally show a neutral face, in accordance with International Civil Aviation Organization (ICAO) requirements for portrait quality.

[0060] The use of photographs such as passport photos has an additional advantage in that facial metrics measured based on horizontal distance in passport photos are not expected to be significantly affected by camera distortions used to collect the passport photo. This is because, relative to the face, a person's eyes and lips are expected to remain on the same vertical axis, regardless of facial expression. Therefore, vertical distortions can be expected to affect the eyes and lips equally. Therefore, vertical distortions are not expected to have a real effect on ratios such as LD / ED, which rely on horizontal distance measurements. On the other hand, horizontal distortions can affect the lip-to-eye ratio metric LD / ED. However, because passport images are generally acquired to International Civil Aviation Organization (ICAO) portrait quality standards, the impact of camera distortions on the metric LD / ED can be expected to be minimal in most cases.

[0061] Therefore, any difference in optical distortion between the camera used to collect the passport photo and the camera used to collect image data while the user is interacting with the automated system is not expected to significantly affect the LD / ED measurements.

[0062] In particular, in practical implementations where a user is interacting with an automated system through their own device, the facial expression detection algorithm may run on hardware with different technical specifications. For example, some older smartphones have lower frame rates than modern smartphones. Thus, in some embodiments, the algorithm is capable of performing facial analysis when running on different hardware with different frame rates. In this way, embodiments of the liveness estimation system intended to run on different types of devices that may be owned by users of the biometric system are "device" agnostic by being frame rate agnostic, provided that a minimum frame rate is available.

[0063] In the context of air travel, this is useful when passengers self-check in on their mobile devices. Self-check-in may occur on a mobile application installed on the mobile device, or through a web-based application that the user can access using a browser on the mobile device. For example, the web application may be supported by a server performing 1:N biometric matching to verify the passenger's identity.

[0064] Consider a case where a slow-speed camera collects images at a frame rate of 10 frames per second (fps) and a high-speed camera collects images at a frame rate of 30 fps. Successive frames captured by the high-speed camera will be captured at time sample points that are only one-third the temporal separation between successive frames captured by the slow-speed camera. Therefore, assuming identical changes in real-time facial expression over a period of time, the differences between successive image samples captured by the slow-speed camera are expected to be greater than the differences between successive image samples captured by the high-speed camera. This can lead to differences between analysis results. For example, if a MACD analysis is performed, an analysis performed on a series of metrics calculated from successive images captured by the slow-speed camera will exhibit higher "momentum" than a MACD analysis result from a series of metrics calculated from successive images captured by the high-speed camera.

[0065] To mitigate the problem, possibly for all embodiments, the facial expression analysis algorithm performs an analysis in which the samples used are at a "target sampling rate." The algorithm may be configured to require a minimum or predefined number of samples (M samples) at the "target sampling rate" to be available. The number "M" may be determined as the number of samples expected over a predetermined time period at the target sampling rate. The M samples are used as a data series for the facial expression analysis. The first time point for the M samples is set to coincide in time with one of the input image frames, and the facial feature metric calculated from that input image frame will be used as the first of the M samples. The image frame that yields the first sample in the M sample series may be the very first input image frame collected. Alternatively, it may be an image frame acquired a predefined time period after the first image frame, or it may be the first image captured once the algorithm determines that the image of the face fits the "target" area in the display view, or it may be an input image frame used as an anchor image.

[0066] For each of the subsequent M-1 samples, the sample value will be the facial feature metric calculated from the input image frame that temporally coincides with the sample, if available. If no input image frame is available that temporally coincides with the required sample, the value for that sample will be determined from the facial metric value calculated from the input image frame that is closest in time to the sample. For example, the sample value may be determined by interpolation between the facial feature metrics calculated from the input image frames that temporally immediately precede and follow the time of the sample.

[0067] The target sampling rate may be a sampling rate expected to be met or exceeded by the frame rates of most camera hardware included in user devices (e.g., most available smartphone or tablet cameras). The facial expression analysis algorithm may be configured to require input images to be collected at or above the minimum actual sampling rate required to generate useful input data for the facial expression analysis algorithm at the target sampling rate.

[0068] As an example, Figure 6 shows an example of how a facial expression analysis algorithm may acquire a data series containing data samples at M separate time points over a 3 second period, as defined by the analysis sampling rate. The sampling rates provided below are merely examples to help explain how the analysis algorithm works, and are not an essential feature of the present invention.

[0069] Figure 6(1) shows the real-time sequential motion that occurs when a person changes their facial expression from neutral to smiling. The images in Figure 6(1) are artificial intelligence-generated images provided for illustrative purposes only and are not actual images collected by a user. The neutral expression is shown in the input image (shown by the image on the left side of Figure 6(1)) that is selected as the "anchor image." The smiling expression can be assumed to be shown in the input image (shown by the image on the right side of Figure 6(1)) captured by a camera a preset time period after the anchor image, e.g., 3 seconds (s).

[0070] In this example, the target sampling rate for the data series to be analyzed is 10 samples per second, as shown in Figure 6(2). However, the actual sampling rate is only 2 samples per second, with input images at times T0, T0+0.5s, T0+1s, etc. If T0 is designated as 0.0 seconds, then facial metrics calculated from the actual images will be at times 0.0 seconds, 0.5 seconds, 1.0 seconds, etc., as shown in Figure 6(3). The system must therefore generate a data series with sample points at the target sampling rate, which in this case means samples are required at times T0, T0+0.1s, T0+0.2s, T0+0.3s, T0+0.4s, etc. It will be understood that the sampling rates mentioned in this paragraph are merely exemplary and should not be construed as limiting how embodiments of the algorithm should be implemented.

[0071] If there are time-matching input frames from the camera for each required sample in the input data series, facial feature metrics calculated from those input frames are used to provide corresponding sample points in the data series, as represented by the dashed arrows between Figures 6(3) and 6(4). Some required sample time points in the data series may not have corresponding input frames from the camera, and therefore may not have metrics or measurements that can be calculated directly from the collected image frames to provide sample values. Therefore, the sample values ​​at each of these time points are estimated from the nearest sample values. The estimation may be interpolation. The interpolation may be linear interpolation.

[0072] The data series can then be used to analyze the motion or distortion characteristics of the observed facial features to estimate whether the face captured in the input image is likely to be a real face or a spoof.

[0073] This method of constructing a data series for analysis has the advantage that it is generally agnostic to variations in the camera's frame rate, at least for cameras that can operate at or above the target frame rate. Additionally, the processing speed of the CPU or processor running the facial expression analysis algorithm is likely to be much higher, at least given current image sensor frame rates. The use of interpolation means that the processing algorithm does not necessarily have to wait for the camera to produce enough frames so that a metric can be calculated to fit the required number of data samples into the data series.

[0074] Facial Motion Analysis The liveness estimation 100 may further include a facial motion analysis algorithm 110 (see FIG. 1 ), which analyzes how the user moves when prompted by the liveness estimation system to move their face, as can be determined from the captured input images.

[0075] Referring to Figure 7, in some embodiments, a user is asked to make movements such that an image of the user's face matches a "target" area on the screen. In some cases, when the target is displayed, it does not remain stationary in the same position on the screen 700, but instead moves or appears in at least one other position, or changes size, or both. The target may therefore be considered a dynamic target. The positioning, sizing, or both of the target 704 may be chosen randomly to make the algorithm more robust against someone attempting to use dynamic spoofs with software that attempts to learn and predict movement patterns.

[0076] 7 and 8, from a "trained reinforcement" perspective, a person interacting with an automated system while facial motion analysis is performed is an "agent." Given that it can be expected that most users will be able to follow instructions, the user may be considered a trained "agent" interacting with the system, where an image 702 of the user's face is shown in a display area 700 and moves within the display area 700 to match the position, size, or both of a target 704. Thus, the display area 700 indicates that the agent will take an action ("A" or " ... t ") "environment," i.e., the "environment" in which the face image 702 is placed on the target 704. The position of the user's face at time t can be calculated by the "state" ("S t ") at time t. t ") thus creates a state S that matches or substantially matches the location of the target 704. t Facial motion analysis determines one or more various factors, such as whether a reward condition is met or the characteristics of the relationship between a change in state and a change in reward over a period of time that the analysis is attempted, to estimate whether a face is likely to be a real face or a spoof.

[0077] FIG. 9 outlines an exemplary implementation of a "Facial Motion" test 900 provided by facial motion analysis. In step 902, the algorithm detects a face in a received image. In step 904, a "target" is displayed at a location on the screen away from the detected image. The target defines an area on the screen. The target's location, and therefore the defined area on the screen, may be randomly selected to become the "randomly selected target." The system instructs the user to perform one or more required movements, so that the user's facial image moves into the area defined by the target. If the algorithm successfully detects a face that fits the target area, the system tracks the detected face across image frames to determine the path it took to move from its starting position (i.e., state "S") to the location in the area defined by the target (step 906). One or more movements, or motions, determined from the image are analyzed (arrow 912).

[0078] The motion analysis may include analyzing the "path" taken by the detected face through the image frames (step 914). The determined path is then analyzed and a liveness estimation is made based on the analysis. This path is represented by arrow 706 in FIG. 7 as the image of the face is aligned with target 704.

[0079] When the facial motion test 900 is performed by a user with the user's real face, the path is expected to be smoother and shorter than the path followed to move a "spoof." For example, some dynamic spoofs use a "brute force" attack, in which a spoof presented to an automated system is moved to random positions by software until it matches the location of a "target." Brute force attacks are therefore likely to result in paths that may not be straight and consistent, even if they successfully align the detected facial image with the "target." Thus, the analysis may be a comparison of the path determined or estimated to have been followed with the path that would be expected if someone were not attempting a spoof attack.

[0080] The analysis of the one or more movements may additionally or alternatively include a determination of the "naturalness" of the one or more movements (step 916). The movements may include only the user's facial movements if the user is interacting with an automated system in which the image sensor is in a fixed position. In embodiments in which the algorithm is intended for applications running on a mobile device, the movements may include facial movements, motion due to camera movements caused by the user, or a combination of both. The order of the path analysis (step 914) and naturalness analysis (step 916) may be reversed from that shown in FIG. 9.

[0081] During the process, if the algorithm does not successfully detect a "real" face that successfully tracks the target, the algorithm may determine that there has been a compliance failure (910) and prompt the user to try again (arrow 908). In Figure 9, the "compliance failure" determination (910) occurs after the face naturalness analysis (step 916), but the "compliance failure" determination may also occur if a failure occurs in one or more of the other steps. For example, a compliance failure may be determined if the system does not detect that there is successful tracking of the face image to the randomly chosen target (failure at step 906), or if the system determines from path analysis that the face image is a spoof image (failure at step 914), or both.

[0082] In video data, the final capture of the video stream (i.e., the last frame or frames) is subject to facial motion, or camera motion if the camera is provided by or as a mobile device, or both. This is because when there is relative motion between the camera and the user's face, this relative motion affects one or more of the position, angle, or depth of the facial features captured by the camera. Relative motion can also affect the lighting or shading that can be observed in the image data. Moving a real face in a three-dimensional (3D) environment, i.e., "natural motion," causes different effects on the observed shading and lighting, as opposed to moving a 2D spoof. Therefore, to estimate the position or motion of the user's face relative to the camera in a physical 3D environment, it is possible to analyze the above-mentioned parameters in the image data and then determine whether the motion is "natural motion."

[0083] The determination of whether a movement is natural, as opposed to a spoofed video playback or brute force presentation attack, may be made by a model trained using machine learning. For example, a reinforcement learning model (FIG. 8) may be used. The user or the user's face may be considered the "agent," and its position in the 3D environment may be considered the "state." The "reward" may be a determination that the movement is natural.

[0084] Based on the path analysis (step 914), the naturalness analysis (step 916), or both, the algorithm makes an assessment, i.e., an estimation, of whether the face is likely to be a spoof (910) or likely to be real (918), meaning there is a compliance failure.

[0085] The facial motion test 900 may be performed several times, i.e., the facial motion analysis may include multiple iterations of the facial motion test 900. The algorithm may require that all of the tests, or a threshold number or percentage of the tests, produce a "normal" result indicating that the face is likely real, before the overall analysis can infer that the detected face is likely an image of a real face.

[0086] Facial dimension analysis The liveness determination 100 may further include a lighting response analysis 112 (see FIG. 1) that analyzes the input image frames to assess discernible responses in the input image frames to changes in lighting.

[0087] In environments where users are interacting with automated systems on mobile devices, some key facial spoofing scenarios may involve using a mobile device screen or printed photograph to match a documented facial image, e.g., a passport image or registration image. The response of a 3D, authentic face is expected to be different from a 2D spoof or a mask worn on someone's face. Therefore, lighting response analysis is sometimes referred to as facial dimension analysis.

[0088] 10 shows an exemplary process 1000 implemented to perform face dimension analysis 112. In the analysis, the input image is checked 1002 to ensure there is a detectable face correctly positioned in the field of view. For the algorithm to work, it is important that there is no significant movement in the detected face and that the presented face is close enough to the screen or camera that the screen brightness, front flash, or both significantly alter the illumination of the face. In some embodiments, light from the mobile screen provides the illumination.

[0089] Once a properly positioned face is detected in the input image, multiple images are captured at different illumination levels. One or more first images may be captured at a first illumination level (1004), and then one or more second images may be captured at a second illumination level different from the first illumination level (1006). Illumination statistics are then calculated for the first and second images to find the occurrence of transitions between "dark" and "light" regions (1008).

[0090] Intensity statistics are calculated for a face region of an image and one or more non-face regions adjacent to the face region. In the example shown in FIG. 11, five regions are defined in the input image, including face region R5 and four other regions R1, R2, R3, and R4 located to the left, right, above, and below face region R5, respectively. Regions R1, R2, and R3 are selected to capture the background of the real person if the person is a real person presenting a real face. Region R4 is a body region that is at a similar "depth" from the screen or camera but is expected to receive slightly less illumination than the face due to the expected positioning of the face relative to the illumination source. Note that FIG. 11 is only an example. Other embodiments may have different regions. For example, another embodiment may not include a "body region" or may include a different number of "background regions."

[0091] Statistics are calculated for the series of input frames captured in steps 1004 and 1006 to determine the variation in intensity contrast between the face region and one or more adjacent regions.

[0092] When the presented face is a "real face," the face will be closer to the screen or camera than the background captured in adjacent regions that are not the user's body. These adjacent regions are expected to be at least a head width behind the face. Therefore, the face region is expected to be better lit than adjacent captured regions of the background. Therefore, when the illumination level is changed, the effect of this change is expected to cause the greatest fluctuation in intensity levels for the face region (e.g., R5 in FIG. 11) compared to the regions where the background behind the person is captured (R1, R2, R3 in FIG. 11).

[0093] For example, referring to FIG. 11 , the algorithm calculates a statistic indicating the luminance contrast between region R5 and one or more of adjacent regions (R1, R2, R3) in the first image, also referred to as the “inter-region” contrast of the first image. The algorithm also calculates a statistic indicating the luminance contrast between those same regions in the second image to obtain the “inter-region” contrast of the second image. The two statistic values ​​are compared to determine the amount of variation in the inter-region contrast (e.g., the difference in intensity of the regions) due to changes in lighting intensity. If there is variation between the inter-region contrast in the first image and the inter-region contrast in the second image, and the amount of variation is above a threshold, then the face is likely to be a genuine face (1014) rather than a spoof (1012).

[0094] Even though exposure compensation impairs absolute changes in luminance levels, relative changes (i.e., contrast) are significantly affected by changes in illumination levels. Therefore, dark-to-light transitions from neighboring regions to face regions are expected to be more pronounced.

[0095] It should be noted that in practice, there may be different ways to implement process 1000. For example, the way to designate a region as "dark" or "light" and the way to calculate the contrast between regions may be implemented by one skilled in the art. One example is applying a threshold to the average pixel intensity, but other methods may be used. Also, the order of processing shown in FIG. 10 may be different. For example, the capture of the second image may occur after or in parallel with the calculation of illumination statistics for the first image.

[0096] Liveness estimation according to different embodiments of the present invention may include one, two, or three of facial expression analysis, facial dimension analysis, and facial motion analysis. Furthermore, where two or more of these analyses are provided, they may be performed sequentially, or, if permitted by the processing power of the hardware used, they may be performed simultaneously. For example, a person may be instructed to make a certain facial expression during facial expression analysis, and illumination levels may be changed so that data collected during that time can also be used to perform facial dimension analysis.

[0097] To meet face to face In one or more of the above-described analyses, the user is asked to position their face so that the image of their face is within a certain target area on the screen. Typically, there is a tolerance for positioning, such that the algorithm considers the image of their face to be "in place" even if there is a slight difference between the area occupied by the image of their face and the target area. Whether the tolerance is large or small, if the user positions their face on the boundary of the tolerance, the user has less freedom to perform further required actions, such as smiling (e.g., in the facial expression analysis example). The user's image of their face may more easily fall outside the tolerance area, and the user may consequently need to repeat the process of repositioning their face.

[0098] To alleviate this problem, the liveness estimation system optionally applies a novel process when fitting the user's facial image to the target area. An example of the face fitting process is shown in FIG. 12. The process begins with a "rough fitting" step 1202 in which the user is required to move so that the contour of the user's facial image is within a first area (1304 in FIG. 13(1)). The first area 1304 is set as a relatively large tolerance range, bounded by dashed lines 1306 and 1308, which represent the lower and upper limits of the first area 1304, respectively. The first area 1304 may also be considered a rough fitting target. The circle 1302 represents the center of the rough fitting target 1304. The dashed line 1306 may also be considered to define a "negative" tolerance from the target center 1302, and the dashed line 1308 may also be considered to define a "positive" tolerance from the target center 1302. In step 1204, after the contour or boundary of the user's facial image is detected to be within the rough fitting target 1304, the position of the boundary of the facial image 1310 (see FIG. 13(2)) is measured. This boundary 1310 is considered to define a reference area. In step 1206, a modified, smaller tolerance range is determined, with the measured boundary 1310 as the center of the smaller range. Referring to FIG. 13(2), the tolerance range 1312 around the measured facial boundary 1310, defined by dashed lines 1314, 1316, is smaller than the rough tolerance range 1304. Steps 1204 and 1206 may both be considered "easy fitting" steps. The resulting tolerance range 1312 may be considered to provide a "easy fitting target."

[0099] In Figures 13(1) to 13(3), the targets and their boundaries are defined by circles, however, they may take other shapes, such as ellipses, shapes resembling the boundary shape of a face, etc.

[0100] The aforementioned coarse and easy fitting processes described with reference to Figures 13(1) and 13(2) reset the coarse fitting target if the image of the user's face falls outside the boundaries of the easy fitting target, and then repeat to reset the easy fitting target (step 1208). The reset coarse fitting area 1320, defined by dashed lines 1322, 1324 that demarcate the target center 1318, is shown in Figure 13(3).

[0101] FIG. 14 schematically illustrates an example of an automated system 1400 for authenticating or registering travelers. The system 1400 includes a device 1402 that includes a camera 1406 configured to collect input image data, or the device 1402 has access to a camera feed. The device 1402 includes a processor 1403, which may be a central processing unit or another processor configuration configured to execute machine instructions to provide all or part of the liveness estimation method described above. For example, the processor 1403 may perform only part of the method if a back-end system using more powerful processing is required to process any of the steps of the liveness estimation method. The machine instructions may be stored in a memory device 1407 co-located with the processor 1403 as shown, or the machine instructions may reside partially or completely in one or more remote memory locations accessible to the processor 1403. The processor 1403 may also have access to data storage 1405 adapted to contain the data to be processed and, in some cases, to at least temporarily store results from the processing.

[0102] The device 1402 further includes an interface arrangement 1404 configured to provide audio and / or video interface capabilities for interacting with the traveler. The interface arrangement 1404 includes a display screen and may further include other components such as a speaker, a microphone, etc. There may also be a communications module 1409, and the device 1402 may wirelessly receive or access data or communicate data or results to a remote location, for example, a separate server computer, a monitoring station computer, or a cloud stage, over a communications network enabling wireless communications 1411.

[0103] In use, input image data is processed by the liveness estimator 1408, which is configured to implement a liveness estimation method. As mentioned above, the liveness estimator 1408 may be provided as a computer program or module, which may be part of an application executed by the processor 1403 of the device 1402. Alternatively, the liveness estimator 1408 may be supported by a remote server or be a cloud-based application, accessible by a web-based application in a browser.

[0104] 14, the box depicting device 1402 is depicted with dashed lines to conceptually indicate that components therein may be provided within the same physical device or housing, or that one or more components may instead be located separately. For example, in embodiments in which device 1402 is a programmable personal device such as a mobile phone or tablet, the mobile phone or tablet may provide a single hardware appliance including an input / output (I / O) interface configuration 1404, a processor 1403, data storage 1405, a communications module 1409, camera hardware 1406, and local memory 1407. Machine instructions for liveness estimation may be stored locally or accessed from the cloud or a remote location, as described above.

[0105] In the illustrated example, the automated system 1400 is used in a travel context. In embodiments where analysis is performed, a passport image 1410 is provided as a reference image for purposes of facial expression analysis performed by the liveness estimator. Providing the passport photo may be by the traveler obtaining a photogram or scanned image of the passport page from the device 1402. In examples where the device 1402 is a kiosk, such as an airport check-in kiosk, the kiosk may include a scanning device configured to scan the relevant passport page.

[0106] In some embodiments, device 1402 is a “local device” that is wirelessly connected to backend system 1412. Such a local device may be provided by a mobile phone or tablet. In the illustrated embodiment, backend system 1412 is a remote server or server system on which 1:N biometric matching engine 1414 resides. Communication between device 1402 and backend system 1412 is represented by dashed double arrow 1411 and may be over a wireless network, such as, but not limited to, a 3G, 4G, or 5G data network, or over a WiFi network. However, backend system 1412 may instead be provided by another server or server system, such as a server at an airport, separate from but in communication with the server performing the 1:N matching.

[0107] In these embodiments, the backend system 1412 may include a backend liveness estimator 1416 configured to implement some or all of the same methods as those implemented by the liveness estimator 1408. In this case, the camera feed data and passport image 1410 are also sent to the backend liveness estimator 1416. That is, while the device's liveness estimator 1408 processes the live camera feed, the camera data is also provided to the backend server 1412 for the same processing. This serves the purpose of performing a validation run of the process to ensure that the results returned by the liveness estimation are not corrupted, or performing steps of the liveness estimation method that may be too computationally intensive for the local device 1402 to process, or both. Only if the results from both the local liveness estimator 1408 and the backend liveness estimator 1416 indicate the "liveness" of the facial image in the camera feed can the automated process of authenticating or enrolling the traveler proceed, and 1:N matching can occur.

[0108] The above describes liveness estimation as part of a check before biometric identification occurs. However, the implementation of biometric identification does not affect the operation of liveness estimation and is therefore not considered part of the present invention in any of the disclosed aspects. For example, liveness estimation may be implemented in a system that does not perform biometric identification. For example, liveness estimation may be implemented in a system that checks whether someone passing through or at a checkpoint is using a spoofed device to hide their identity or to pretend to be someone else, for example, using a "spoof" to join a video conference or to register themselves in a specific user database.

[0109] Variations and modifications may be made to the previously described portions without departing from the spirit or scope of the present disclosure.

[0110] In the following claims and the foregoing description, unless the context requires otherwise by express language or necessary implication, the word "comprises" or variations such as "comprises" or "comprising" is used in an inclusive sense, i.e., to specify the presence of stated features, but does not exclude the presence or addition of further features in various embodiments of the disclosure. [Explanation of symbols]

[0111] 700 screens 702 Face Images 704 Goal 900 Facial Motion Test 1400 Automated System 1402 devices 1403 processor 1404 Input / Output (I / O) Interface Configuration 1405 Data Storage 1406 Camera 1407 Memory Devices 1408 Liveness Estimator 1409 Communication Module 1410 Passport Images 1411 Wireless Communications 1412 Backend system, backend server 1414 1:N Biometric Matching Engine 1416 Backend Liveness Estimator

Claims

1. 1. A method of estimating whether a presented face of a user is a genuine face by analyzing collected image frames of the presented face, the method comprising: (a) determining and analyzing one or more movements made by the user to align a facial image of the presented face with a randomly selected target; (b) determining and analyzing changes in facial feature metrics in the facial image in response to changes in the user's facial expression; and (c) determining and analyzing the effect of illumination variations on the contrast level between a facial region and one or more regions adjacent to the facial region in the collected images; The method includes one or more of:

2. 10. The method of claim 1, wherein (a), (b), and (c) are performed sequentially.

3. 10. The method of claim 1, wherein at least one of (a), (b), and (c) is performed simultaneously with another one of (a), (b), and (c).

4. determining and analyzing movements made by the user to align a facial image of the presented face with a randomly selected target; displaying a randomly selected target on a display screen and instructing the user to perform a motion on the display screen to move the user's facial image from its current position so that the user's facial image tracks the randomly selected target; analyzing image frames collected while the user is instructed to track the randomly selected target; estimating whether the presented face is a genuine face based on the analysis; and 4. The method of claim 1, comprising:

5. 5. The method of claim 4, wherein analyzing the image frames collected while the user is instructed to track the randomly selected target includes determining one or more movements made by the user.

6. 6. The method of claim 5, wherein estimating whether the presented face is a genuine face comprises comparing the determined one or more movements with movements that a person would expect to make.

7. The method of claim 6 , wherein the estimation is performed by a machine learning model.

8. The method of claim 7 , wherein the estimation is performed by a reinforcement learning base trained using data on natural movements.

9. 9. The method of claim 5, wherein determining a movement comprises determining a path between the current location and the randomly selected target.

10. determining and analyzing changes in facial feature metrics in the facial image in response to changes in the user's facial expression, instructing the user to make facial expressions that cause distortion of visually detectable facial features; analyzing image frames collected while the user is instructed to make the facial expression; estimating whether the presented face is a genuine face based on the analysis; and 10. The method of any one of claims 1 to 9, comprising:

11. analyzing the image frames collected while the user is instructed to make the facial expression, obtaining a time series of metric values ​​from the series of analyzed frames by calculating a metric based on the location of one or more facial features from each analyzed frame; 11. The method of claim 10, comprising:

12. 12. The method of claim 11 , wherein estimating whether the presented face is a genuine face based on the analysis comprises determining momentum in the time series of metric values ​​and comparing the momentum to an expected momentum profile for a genuine, smiling face.

13. 13. The method of claim 10, wherein the user is instructed to smile.

14. 14. The method of claim 13, wherein for each analyzed frame, the calculated metric is or includes a ratio of the distance between the eyes of a detected face in the analyzed frame to the width of the mouth of the detected face.

15. 15. The method of claim 10, further comprising the steps of: comparing the image frames collected while the user is instructed to make the facial expression with reference images in which the user has a neutral expression; and selecting from the analyzed image frames the image that is most similar to the reference image as an anchor image in which the user is considered to have a neutral expression.

16. analyzing the effect of changes in illumination on the presented face on the image data, capturing one or more first image frames of the presented face at a first illumination level and capturing one or more second image frames of the presented face at a second illumination level different from the first illumination level; analyzing the first image frame to determine a first contrast level between a detected face region in the first image frame and an adjacent region adjacent to the detected face region; analyzing the second image frame to determine a second contrast level between the detected face region in the second image frame and an adjacent region adjacent to the detected face region of the second image frame, wherein the relationship between the adjacent region and the detected face region of the second image frame is the same as the relationship between the adjacent region and the detected face region of the first image frame; comparing the first and second contrast levels to estimate whether the presented face is likely to be a genuine face; 16. The method of any one of claims 1 to 15, comprising:

17. 17. The method of claim 16, wherein comparing the first and second contrast levels comprises determining whether a change between the first contrast level and the second contrast level is greater than a threshold value.

18. During the step in which the user is requested to align the user's facial image with a target area on a screen, the method further comprises: prompting the user to align a facial image of the presented face of the user with a larger first target area; Upon detecting a face image within the first target area, setting an area generally encompassed by the face image as a reference target area.

18. The method of any one of claims 1 to 17, comprising:

19. 20. The method of claim 18, further comprising applying a tolerance range around the larger first target area, whereby detection of a facial image within the tolerance range triggers setting the area of ​​the facial image as the reference target area.

20. 20. The method of claim 18 or 19, comprising applying a tolerance range around the reference target area, wherein the facial image of the presented face of the user is considered to remain within the reference target area if it is within the tolerance range around the reference area.

21. 21. An apparatus for estimating whether a presented face of a user is a genuine face by analyzing collected image frames of the presented face, the apparatus comprising: a processor configured to execute machine instructions that implement the method of any one of claims 1 to 20.

22. The apparatus of claim 21 , wherein the apparatus is a local device used or accessed by the user.

23. The apparatus of claim 22 , wherein the local device is a mobile phone or a tablet.

24. 24. The apparatus of claim 21, wherein the collected image frames are sent to a backend system via a communications network and processed by a processor of the backend system configured to execute machine instructions that at least partially implement the method of any one of claims 1 to 20.

25. 25. The device of claim 24, wherein the presented face is presumed to be a genuine face if the processing results by the processor of the device and the processing results by the processor of the backend system both presume the presented face is a genuine face.

26. 26. The device of claim 25, wherein the device allows the user to interface with an automated system biometric matching system that enrolls or verifies the user's identity.

27. 27. The apparatus of claim 25 or 26, wherein the user is a passenger on an air trip and the backend system is a server system that hosts a biometric matching service or is a server system connected to another server system that hosts a biometric matching service.

28. 1. A method for biometrically determining a subject's identity, comprising:

21. The method of claim 1, further comprising: estimating whether the presented face of the subject is a genuine face; If the presented face is presumed to be a genuine face, providing the collected two-dimensional image of the presented face for biometric identification of the subject; outputting the result of the biometric identification; A method comprising:

29. 30. The method of claim 28 performed during the check-in process by an air travel passenger.