Method for securing access to a data source
A biometric authentication method using a time-varying light signal on terminals addresses fraud vulnerabilities by evaluating the contribution and noise level of acquired video signals, ensuring secure and accessible remote access.
Patent Information
- Application Number
- FR2024008186
- Authority / Receiving Office
- FR · FR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-24
- Publication Date
- 2025-11-07
AI Technical Summary
Existing biometric authentication methods for remote access are vulnerable to frauds such as injection and presentation attacks, particularly in mobile applications, which can be inconvenient for users with disabilities and lack robustness against fraud.
A method using a terminal with a light source and camera that applies a time-varying, periodic signal for biometric authentication, evaluating the presence of fraud by estimating the contribution of the emitted signal and comparing it to the noise level, ensuring the acquired video was captured in real-time and not from a third-party object.
This method provides robust protection against fraud, is accessible to all users, including those with disabilities, and does not require complex image analysis, ensuring secure access to data sources.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Title of the invention: Method for securing access to a data source. Technological background.
[0001] The present invention relates to the field of securing remote access to data sources, whether during enrollment (for example, to create an online account) or during account access. Indeed, mobile applications typically require biometric authentication of the user on their connected device (such as a mobile phone, smartwatch, tablet, or computer) for remote access. To this end, a camera on the connected device remotely acquires an image of the user's face in order to authenticate the user.Nevertheless, frauds exist during this type of acquisition, so it is necessary to secure this step and detect potential frauds, in particular injection fraud (emulation of a virtual camera) (called in English: "Injection Attack") or presentation fraud (called in English: "Presentation Attack"), which notably includes frauds shown on paper, screen.
[0002] It is known, for example in document FR3100074, to ask the user to perform a randomly determined gesture, but this can be inconvenient for the user, particularly if the latter suffers from a disability affecting their mobility. Presentation of the invention
[0003] The invention aims to remedy at least some of these drawbacks and preferably all of them, and in particular aims to provide a method of securing remote access to a data source which is robust to fraud by injection or presentation and accessible to all, without discomfort or effort on the part of the user.
[0004] According to one aspect of the invention, a method is proposed for securing remote access to a data source from a terminal by a user of said terminal, the terminal comprising a light source and a camera directed towards said user, said method comprising steps of: - acquisition of a time-lapse video signal by the camera; - control, during the acquisition stage, of the light source according to a time-varying, periodic signal of which at least one parameter is time-invariant and of randomly fixed value; - evaluation, based on an initial series of images of the acquired time-domain signal, of the presence of fraud by injection or presentation by: - estimation of a contribution of the emitted signal imposed on the acquired video time signal as a function of said parameter - evaluation of the noise level of the acquired time-domain video signal; - comparison of the contribution in relation to the noise level; - authorization to continue access to the data source depending on the result of the comparison step.
[0005] This method addresses the drawbacks listed above and is preferably applied to a connected mobile terminal (such as a mobile phone, smartwatch, tablet, or computer). It thus secures access both in an enrollment context (for example, for creating a bank account) and in a consultation context (for example, for proof of identity in an administrative procedure or for accessing online accounts). The method can also be applied to a fixed terminal (such as an airport kiosk or a physical access system, where the remote data source is, for example, a list of authorized access persons).Indeed, the random (including pseudo-random) nature of the signal presents a challenge and allows verification that the acquired video was captured by the terminal's camera and shows the user facing the camera in real time at the moment of acquisition, and not from a third-party object or recorded at a different time. This process, particularly through the analysis of the acquired video signal, which includes information in three dimensions—height, image width, and time—for one (or more) channel (per color component), ensures that the acquired video was indeed captured by the terminal's camera and provides robust protection against screen fraud or the presentation of printed photos, for example.Since the parameter and its value are known at the evaluation stage, the process does not require estimation of the challenge parameter in the acquired signal, nor any specific synchronization with a dedicated event to identify a specific sequence. Therefore, there is no need to analyze and compare the acquired and expected images one by one, making the process more flexible and less memory-intensive. Finally, if the score indicates the presence of fraud, access to the data source cannot continue.
[0006] Advantageously, the method may include an intermediate step of determining the probability of fraud by injection or presentation as a function of the calculated signal-to-noise ratio; with authorization to continue access to the data source if the probability of fraud by injection or presentation is less than a first predetermined threshold.
[0007] Equivalently, a probability of non-fraud can be determined by injection or presentation and in this case the threshold condition applies if said probability is greater than another predetermined first threshold.
[0008] Advantageously, the contribution is expressed as a value, unique (global) or vectorial (also called matrix), time-dependent (such as in the form of a signal) or not (such as in the form of energy).
[0009] In one embodiment, the imposed time-varying emitted signal is a periodic color component whose time-invariant parameter is a fundamental frequency, thus allowing for simple implementation of the imposed signal and a simple and rapid spectral analysis of the acquired signal. Furthermore, for a given imposed periodic signal, which may, for example, have several color components, these components can be differentiated, with different fundamental frequencies.
[0010] In one embodiment, the imposed time-varying signal comprises at least two periodic color components of the same fundamental frequency or multiples thereof, with a phase shift between said at least two components, said at least one time-invariant parameter of the imposed signal being said fundamental frequency or the phase shift. This allows, in particular, for several possible invariant parameters or for combining them: a fundamental frequency, for example, of a single and identical value for each time-varying color component of the imposed signal, and a phase shift between said components, which makes the challenge more diverse.
[0011] Preferably, in the visible spectrum, only the red and blue components are imposed, which facilitates analysis because this choice makes the process robust to color crosstalk, linked for example to the fact that the red channel of the camera sees a little of the neighboring green channel of the light source, the green component remaining invariant over time.
[0012] Advantageously, the imposed time-varying signal is applied to an infrared component, which makes it possible to implement the light variation process outside of visible wavelengths without the user noticing.
[0013] In one embodiment, the imposed fundamental frequency is set randomly within a working range between 0.5Hz and 3Hz, which promotes user comfort and allows for the use of a standard sampling of 15 frames per second, for example, the fundamental frequency being unique for all the components concerned or differentiated by component.
[0014] Advantageously, a determination of the value of each randomly fixed invariant parameter is implemented by randomly drawing from a set of values of the working range relative to said invariant parameter or by pseudo-random draw from said set of values from which are excluded the values previously determined from said invariant parameter for said same terminal and / or user.
[0015] Advantageously, the said values previously determined for said same terminal and / or user are stored in an exclusion register in association with a terminal identifier and / or a biometric user identifier to which they have been applied.
[0016] In one embodiment, the terminal comprises a display screen, the light source being all or part of the terminal's display screen, and the control of the light source corresponding to a display on all or part of the screen with a color varying according to the time-varying signal imposed. This preferred method allows the terminal to be held in the usual way, particularly if the terminal is a tablet or a mobile phone, without having to, for example, turn it over, thus avoiding handling difficulties and remaining discreet.
[0017] Advantageously, at least one imposed time-varying color component is of the square wave type, which allows a simple signal to be driven.
[0018] Advantageously, the imposed time-variable signal is sinusoidal, that is to say that the component or at least one of the imposed time-variable color components is sinusoidal, which is simple to control, and in particular if both time-variable color components of the signal are sinusoidal, the spectral analysis is further simplified because only one peak appears per color component.
[0019] Advantageously, at least one imposed time-varying color component is of the triangular type, which allows for easy control.
[0020] Advantageously, the imposed time-varying color components are of different types, with, for example, a triangular blue component and a sinusoidal red component, this mixture allowing the diversity of the challenge to increase.
[0021] In one embodiment, the contribution estimation includes bandpass time-domain filtering or a predetermined component selection of a fast Fourier transform of the acquired video time-domain signal, or a correlation with the imposed emitted signal. This makes it particularly easy to isolate the contribution of the imposed emitted signal to the acquired video time-domain signal since the invariant parameter is known.
[0022] In one embodiment, the contribution is expressed as energy, the evaluation of the noise level of the acquired time-domain video signal being implemented by calculating the energy of a residual signal corresponding to the subtraction of a time average of the acquired time-domain video signal from said acquired time-domain video signal, which makes it possible to integrate the temporal dimension and define a level of noise from the acquired time-domain video signal simplifies subsequent comparisons. The noise level thus evaluated by calculating the residual signal energy can be obtained by calculating the energy of the signal resulting from subtracting a time average of the acquired time-domain video signal from said acquired time-domain video signal, or alternatively, equivalently in two steps: by applying a bandpass filter corresponding to the working range to the acquired signal, then calculating the energy of the acquired signal thus filtered, from which the energy of the contribution is subtracted, the result being the noise level of said acquired time-domain video signal.
[0023] According to one embodiment, the step of comparing the contribution with respect to the noise level includes a calculation of a signal-to-noise ratio by dividing the contribution by the noise level, which then allows the step of authorizing continued access to the data source to authorize access based on the calculated signal-to-noise ratio.
[0024] In one embodiment, at the evaluation stage, based on a first series of images of the acquired time signal, the presence of fraud by injection or presentation: - at least one area of interest is determined, in each image of the first series, said area of interest being common to the images of the first series, (that is to say imaging the same surface of the acquired object, for example a person's face, on said images of the first series) -a temporal signal observed representative of the pixel values of the images of the first series of images over all or part of the area of interest is determined; - During estimation, the temporal signal of the first series of images is the observed temporal signal determined by region of interest. This allows the steps of the process to be discretized by regions of interest by analyzing one signal per region of interest (by applying the process in parallel to each region of interest or by applying it to a signal vector) in order to reduce the computational load on the processor responsible for executing the computational step, or to facilitate computation on parallelized processors. Another advantage is the ability to focus on a single global area (for example, the upper part of the face) or on regions of interest significant for fraud detection.
[0025] In one embodiment, the step of - evaluation of the noise level of the acquired time-domain video signal; Or - comparison of the contribution relative to the noise level, preferably by calculating the signal-to-noise ratio by dividing the contribution by the noise level, includes an average calculation, notably weighted by area of interest, or of median of said areas of interest so that the calculated signal-to-noise ratio is global.
[0026] In one embodiment, the process comprises: - an evaluation step, based on a second set of images from the acquisition, to detect the presence of fraud by presentation by: - determination of at least one characteristic surface, in each image of the second series, said characteristic surface being common to the images of the second series (i.e. imaging the same surface of the acquired object, for example a person's face, on said images of the second series), - determination of a temporal signal observed by a characteristic surface, - estimation, using a characteristic surface, of the contribution of the imposed signal to the observed signal in order to form a feedback vector, - A score is determined by applying a classifier to an audit vector, and the authorization step for continued access to the data source is based on the determined score. This additional verification, by analyzing the response of the captured object to the temporally variable lighting emitted by the controlled light source, secures data access. This additional verification provides robustness against screen fraud or fraud through the presentation of printed photos, for example, and allows for the consideration of one or more dedicated characteristic surfaces significant from the perspective of the expected reflection in relation to the user's expected local shape within said characteristic surface. It also reduces the size of the data to be transmitted in the case of a contribution estimation step, which is performed as before, and / or a score determination step that is performed remotely on a server, for example.
[0027] In one embodiment, the area of interest or characteristic surface belongs to a detected region of interest imaging all or part of a face, supposedly that of the terminal user, which does not require any specific manipulation by the terminal user, the latter being accustomed to orienting it towards his face, the front camera performing the acquisition.
[0028] Advantageously, prior to the step of determining the area of interest or the characteristic surface, a spatial alignment step is implemented for each image in the series, of said images of the first or respectively second series with respect to a reference point of the detected region of interest, for example by eye registration, which makes it easier to determine the area of interest or characteristic surface corresponding for example to the same portion of the face in the different images of the first or respectively second series.
[0029] Advantageously, the area of interest or characteristic surface is a triangle of a mesh of the region of interest, or a pixel in the vicinity of an interest point of the region of interest, said region of interest or characteristic surface being present in all images of the first or second series of images respectively. For example, a point of interest is a semantic point of the face.
[0030] Advantageously, the area of interest and the characteristic surface are distinct, which allows the analysis to be focused on distinct areas of the face depending on the type of fraud analysis performed.
[0031] Alternatively, the area of interest and the characteristic surface are identical so as to be able to share the steps of detecting the latter.
[0032] Advantageously, said classifier is implemented by a neural network, which allows for fast processing with good classification performance.
[0033] Advantageously, the method may include an intermediate step of determining the probability of fraud per presentation as a function of the determined score; with authorization to continue access to the data source if the probability of fraud per presentation is less than a second predetermined threshold, which allows in particular a normalization of the calculated ratio allowing then a comparison of the probability obtained at said second predetermined threshold, regardless of the type of invariant parameter.
[0034] Equivalently, a probability of no fraud per presentation can be determined and in this case the threshold condition applies if said probability is greater than another second predetermined threshold.
[0035] Advantageously, if several time-varying time signals are imposed, as many time signals are observed per characteristic surface, the return vector being formed of as many sub-vectors as there are time-varying signals imposed, which makes it possible to have at the input of the classifier a vector of vectors (also called a matrix, with for example as many columns as there are imposed signals (periodic color components in particular) and as many rows as there are characteristic surfaces) making the classifier robust.
[0036] Advantageously, the method may include a step of determining, by characteristic surface, the contribution of other lighting sources by subtracting the contribution of the imposed signal from the observed variable signal in order to form a subvector of the feedback vector. This allows the use of a vector of vectors as input to the classifier, thus improving its performance. Furthermore, since the contribution of the imposed signal to the acquired signal is small, the contribution of other lighting sources can also be approximated by a simple time average of the acquired signal.
[0037] In one embodiment, the method includes a step of evaluating, from a third series of images from the acquisition, the presence of specific fraud by recapture or injection by: - Determination of an observed latency based on the fundamental frequency of the imposed periodic time-varying signal and an estimated phase shift between the imposed periodic time signal and an observed time signal representative of the pixel values of the images in the third image series over all or part of the detected region of interest common to said images in the third image series; - Calculation of a difference between the observed latency and a reference latency previously obtained for said terminal or for the same terminal model; and authorization to continue access to the data source being a function of the calculated latency difference. This additional step makes it possible to verify the consistency between the observed latency and the reference latency and thus to better identify fraud by recapture or injection.
[0038] Advantageously, the method may include an intermediate step of determining the specific probability of fraud by recapture or injection as a function of the calculated latency difference, with authorization to continue access to the data source if the specific probability of fraud by recapture or injection is less than a third predetermined threshold, which notably allows a normalization of the calculated latency difference allowing then a comparison of the specific probability to said third predetermined threshold.
[0039] Equivalently, a specific probability of non-fraud by recapture or injection can be determined and in this case the threshold condition applies if said probability is greater than another predetermined third threshold.
[0040] Advantageously, a camera acquisition parameter, in particular an image size, is modified during acquisition with an increase in the specific probability of fraud by recapture or injection in the absence of a change in the latency observed during acquisition.
[0041] Advantageously, the second or third series of images of the acquisition and the first series of images of the acquisition are identical.
[0042] According to another aspect of the invention, a computer program is proposed comprising instructions adapted to the implementation of each of the steps of the process according to the invention when said program is executed on a computer.
[0043] According to another aspect of the invention, a non-transient information storage means is proposed, removable or not, partially or totally readable by a computer or a microprocessor comprising code instructions of a computer program for the execution of each of the steps of the process according to the invention.
[0044] Advantageously, a device according to the invention comprises said terminal and an information processing unit implementing said process, the processing unit comprising said non-transient information storage means, the processing unit being able to be located in the terminal or distributed between the terminal and a remote server, which allows for a device with a secure operating architecture. Presentation of the figures
[0045] The invention will be better understood from the following description, which relates to embodiments and variants of the present invention, given by way of non-limiting examples and explained with reference to the accompanying schematic drawings, in which:
[0046] [Fig-1] [Fig.1] schematically shows a person presenting their face to a terminal during the implementation of the security process according to one possible embodiment of the invention,
[0047] [Fig.2] [Fig.2] represents a schematic block diagram of a unit of information processing of a device capable of implementing one or more embodiments of the invention,
[0048] [Fig.3] [Fig.3] shows a schematic diagram of the steps implemented in the a security method, according to a possible embodiment of the invention,
[0049] [Fig.4] [Fig.4] shows a schematic diagram of complementary steps put implemented in the securing process, according to a possible embodiment of the invention, and [Fig.5] [Fig.5] shows a schematic diagram of additional steps implemented in the securing process, according to a possible embodiment of the invention.
[0050] Identical references will be used from one figure to another to designate identical or similar elements, in their form or in their function.
[0051] For the sake of brevity, the term substantially refers to values to the extent of plus or minus 10%. Detailed description
[0052] The invention can be applied in various contexts. The illustrated embodiment relates to securing remote access to a data source from a user terminal, here a smartphone, by a user of the mobile terminal. The smartphone includes a light source, here a portion of the phone's screen, and a front-facing camera oriented towards the user to acquire images of their face. Another embodiment (not illustrated here) of the invention relates to acquiring another biometric characteristic, such as a dermatoglyph of the mobile phone user, for example, by the (main) rear camera of the mobile phone, the light source being, for example, the flash of said rear camera.In all cases, the aim is to detect fraud by manipulating the light signal in such a way that its variations have a time-invariant property, fixed randomly.
[0053] The method according to the invention can be used in various applications. In particular, the invention can be used to implement a method for monitoring a driver, or to implement fraud detection within the framework of biometric authentication for enrollment or data access. In all cases, an estimation of the contribution of the imposed signal to an acquired time-domain signal is performed, knowing the invariant parameter, and a noise level of the acquired video time-domain signal is used to determine the presence or absence of fraud.
[0054] For the sake of simplicity and in an illustrative and non-limiting manner, the invention will be presented below in the context of a biometric method for authenticating a face, but the principles can be applied to any application involving facial recognition. In this context, to verify the authenticity of the presented face, the processing unit implements a fraud detection method by calculating a signal-to-noise ratio by dividing the contribution by the noise level.
[0055] Latency time corresponds to the response time of the entire acquisition chain, namely the terminal, and is broken down in particular into a delay for illumination by the light source (in particular display by the screen) and an acquisition delay from the camera.
[0056] With reference to [Fig. 1], the authentication method can be implemented using a device comprising a terminal 1 to which a user's face 2 is presented. The terminal 1 includes an information processing unit and a camera 3 adapted to acquire image streams of objects presented in its acquisition field 4. Preferably, the terminal 1 also includes a screen 5 capable of displaying images to the user, and is configured so that the user can simultaneously present their face 2 in the acquisition field 4 of the camera 3 and view the screen 5. The terminal 1 can thus be, for example, a mobile terminal such as a mobile phone (in particular a "smartphone"), smartwatch, or tablet, which typically has an ideal configuration with a large screen 5 and a light weight that allows for handling in selfie mode similar to a mobile phone.Terminal 1 can nevertheless be any type of computerized device, and in particular can be a computer with a camera or a fixed kiosk dedicated to identity checks, for example installed in an airport. Terminal 1 can also be an electronic subsystem embedded in a vehicle forming a connected system for driver recognition, or for accessing applications for the driver or passenger. The information processing unit comprises at least a processor and memory, and allows the execution of a computer program for implementing the method according to the invention.
[0057] Figure 2 is an example of a schematic block diagram of an information processing unit 106 for implementing one or more embodiments of the invention. The information processing unit 106 typically comprises at least one calculator, computer, microprocessor, or other device enabling the execution of a computer program responsible for controlling the various stages of the process according to the invention. The information processing unit 106 includes a communication bus connected to: - a central processing unit 601, such as a microprocessor, denoted CPU; - a 602 transient memory, noted as RAM, to store the executable code of the process of implementing the invention as well as the registers adapted to record variables and parameters necessary for the implementation of the process according to embodiments of the invention; the memory capacity of the device can be supplemented by an optional RAM memory connected to an expansion port, for example; - a non-transient memory 603, denoted FLASH, for storing computer programs and calibration data for the implementation of the embodiments of the invention; the stored computer programs include in particular a computer program comprising instructions adapted to the implementation of each of the steps of the process according to the invention when said program is executed on the processing unit 106, said FLASH memory 603 is then an example of a non-transient means of storing information, removable or not, partially or totally readable by a computer or a microprocessor comprising code instructions of the computer program for the execution of each of the steps of the process according to the invention; - A 604 network interface, denoted NET, is normally connected to a communication network on which digital data to be processed is transmitted or received. The 604 network interface can be a single network interface or composed of a set of different network interfaces (e.g., wired and wireless, or different types of wired or wireless interfaces). Data packets are sent over the network interface for transmission or are read from the network interface for reception under the control of the software application running in the 601 processor. - A 605 GUI user interface for receiving input from a user or for displaying information to a user, including guidance information (voice and / or visual), including, in this case, screen 5. - an input / output module 607 for receiving / sending data to / from external devices such as a hard drive, removable storage media, or others.
[0058] The executable code can be stored in non-volatile memory 603, for example, flash memory or read-only memory, or on removable digital media such as such as a disk. According to one variant, the executable code of programs can be received by means of a communication network, via the network interface 604, in order to be stored in one of the storage means of the information processing unit 106, such as the FLASH memory 603, before being executed.
[0059] The central processing unit 601 is adapted to command and direct the execution of instructions or portions of software code of the program or programs according to one of the embodiments of the invention, instructions which are stored in one of the aforementioned storage means, such as the FLASH memory 603. After power-up, the CPU 601 is capable of executing instructions from the transient RAM memory 602, relating to a software application. Such software, when executed by the processor 601, causes the execution of the method according to the invention.
[0060] In this embodiment, the device is a programmable device that uses software to implement the invention. However, as a subsidiary measure, the present invention can be implemented in hardware (for example, in the form of a specific integrated circuit or ASIC (application-specific integrated circuit) or in the form of a programmable logic component or FPGA (field-programmable gate array).
[0061] The information processing unit 106, as illustrated, is local to the terminal 1. This is, for example, the preferred architecture in the case of a fixed terminal, such as a fixed kiosk dedicated to identity checks, because these kiosks are secured without user access, and it is then recapture fraud that is critical and monitored by the method according to the invention. Alternatively, the information processing unit 106 may be external to the terminal 1, or distributed and comprise multiple processing subunits, including at least some external to the terminal 1 (particularly in the case of a mobile terminal as previously illustrated), communicating with each other via the network interface. Similarly, depending on the nature of the terminal, all or part of the memory may be physically remote, hosted, for example, on a remote server.For example, in a fixed terminal 1, the terminal acts as the master, and the initialization, acquisition, and control modules are hosted locally within the user terminal 1. However, other modules may not be, or only partially, hosted locally, but rather in a physically remote processing entity (slave), such as a remote server. This sharing of computations between the local user terminal 1 and the remote server allows only the information necessary for decision-making to be sent to the remote server, thus minimizing response time related to data exchange and network throughput, without compromising client-side security related to reverse engineering. Terminal 1 can, for example, exchange acquired biometric information with the remote server, as illustrated in the document. FR1661737. This may include having redundant calculations, with the remote server verifying all or part of what terminal 1 has done. Alternatively, in particular for a mobile terminal 1, but applicable to any type of terminal, the remote server is the master and the user terminal 1 the slave, so that the invariant parameter is defined and its random draw executed by the remote server, then the imposed signal transmitted to the user terminal so that the latter implements the control accordingly, which maximizes the security of the user terminal 1 and prevents a replay on the user terminal side, the latter not unilaterally deciding the challenge and being agnostic of the chosen invariant parameter and / or its value when the control instruction is advantageously sent in real time by the server to the user terminal.Similarly, the user terminal 1 can then send the raw acquired signals directly to the remote server to minimize local processing and reduce the risks of reverse engineering on the client side, or conversely, send the information directly necessary for decision-making (for example, the signal-to-noise ratio) to minimize network load and response time related to data exchange, which is dependent on network bandwidth. Furthermore, the exchanged information, particularly from the server to terminal 1, can be encrypted to improve the security of the exchanges, especially during the transmission of the imposed signal.
[0062] With reference to [Fig. 3], the user of terminal 1 uses a device according to the invention to authenticate themselves to an online account via an application. A biometric facial authentication process is then initiated, prompting the user to present their face 2 in the acquisition field 4 of the camera 3. Before generating the biometric template and accessing the user's enrolled template for comparison, the biometric authentication process calls the remote access security process P to detect fraud attempts during access to the online account.
[0063] When called by the biometric facial authentication process, the security process P, the information processing unit 106 of the device implements the initialization step E0 of the security process P corresponding in particular to the verification of the operation of the camera 3 of the terminal 1 of the device.
[0064] The security process P then continues with the implementation, by the information processing unit 106, of the instructions to: - determination of the invariant parameters (not shown) by random selection of the value(s) of the invariant parameter(s), i.e. of the fundamental frequency(ies) of the imposed signal and / or of the phase shift between two components of the imposed signal; - acquisition El of a time-lapse video signal by camera 3, the orientation of which towards the user's face was requested from the user; - control E2, during all or part of the acquisition stage, of the light source, here the screen 5, according to the time-varying signal emitted imposed by displaying on more than half of the screen 5 a color varying according to the time-varying signal imposed; The first series of images, as described here, results from a selection of acquired images based on quality criteria (sharpness) and the fact that these images include a face (region of interest) – which is particularly important with regard to the inter-eye distance, which should preferably extend over more than 50 pixels. This selection therefore results from image analysis, notably implemented using a neural network.Furthermore, it should be noted that the first series of images may only contain images from a time window of the acquisition period El by camera 3, preferably the duration of said time window is greater than or equal to twice the maximum period of the imposed signal(s) (taking into account the different imposed components) with a sampling frequency greater than or equal to twice the maximum frequency of the imposed signal(s) (taking into account the different imposed components); - evaluation E3, from a first series of images of the acquired time signal, of the presence of fraud by injection or presentation (on screen or by presentation of printed photo for example) by: . - E3d estimation of a contribution of the emitted signal imposed on the acquired video time signal as a function of the invariant parameters; - E3e evaluation of the noise level of the acquired temporal video signal; - comparison E3f of the contribution relative to the noise level by calculating a signal-to-noise ratio by dividing the contribution by the noise level.
[0065] Authorization E6 to continue access to the data source based on the calculated signal-to-noise ratio. Otherwise, the biometric authentication process is interrupted and an error message may be displayed on screen 5. In both cases, the time-stamped status, success (no fraud) or failure (fraud detected), is preferably recorded in a local RAM register or in an external register.
[0066] It is noted that the steps of estimating a contribution E3d and evaluating a noise level E3e can be reversed or carried out simultaneously.
[0067] In the embodiment illustrated and described in connection with [Fig.3], the evaluation step E3 is applied according to a discretization by areas of interest, so as to describe this embodiment in detail, however this embodiment is not exclusive nor limiting, the method according to the invention also being applied globally on a single area which would be for example the region of interest, corresponding to the area of each image of the first series imaging the face of the user.
[0068] In this embodiment, the image selection can also be refined so that only those images containing all the areas of interest are included in this first series of images. In this case, the alignment steps E3a and the area-of-interest determination steps E3b are performed prior to the image selection and the fraud detection step E3, which are based on the first series of images of the acquired time-domain signal. It should be noted that it would also be possible to define as many first series of images as there are areas of interest.Indeed, if during the acquisition the person moves, showing one profile and then the other, it is possible to determine a first set of images for a region of interest specific to one side of the nose and another first set of images (distinct from the previous images) for a region of interest specific to the other side of the nose, whereas for a region of interest specific to the forehead, for example, the first set of images used could include all the images from these first two sets. In the implementation method illustrated here, only the criteria of quality and presence of a face (with an inter-eye distance greater than 50 pixels) in the image are applied, which means that all the images in the first set contain all the regions of interest, for the sake of simplification.
[0069] The E3 evaluation step, based on the first series of images of the acquired time-domain signal, for the presence of fraud by injection or presentation, comprises: - For each image in the first series, spatial alignment (E3a) of said images with respect to a reference point of the detected face, for example by eye registration; this optional alignment step then allows working in a fixed reference frame and without the effect of scaling over time, even if the person moves closer; It should also be noted that, alternatively, all or part of the pre-processing steps described above can be implemented in conjunction with this alignment step. - Determination (E3b) of at least two areas of interest in the detected face, each area of interest being common to the images in the first series; - E3c determination of an observed time-domain signal per region of interest, representative of the pixel values of the images in the first series of images; here the result is expressed as a vector of signal vectors, also called a matrix, each column corresponding to a component (blue and red) and each row to a region of interest, each element of the matrix comprising the observed time-domain signal vector of the corresponding region of interest spatially averaged over the pixels of said region of interest; - E3d estimation, per region of interest, of a contribution of the imposed emitted signal to the observed time-domain signal as a function of two invariant parameters, whose nature and value are known, the contribution per region of interest being here, but not limited to, expressed as an energy in order to manipulate a contribution estimate in non-time-domain vector form; alternatively, it could have remained in the form of a signal vector or even written as a single value, the contributions of the different zones being averaged for example (with or without the application of weightings), - evaluation E3e of a noise level of the acquired time-domain video signal by determination, by area of interest, by calculation of an energy of a residual signal corresponding to a subtraction of a time average of the acquired time-domain video signal from said acquired time-domain video signal; this calculation of the energy of a residual signal is here carried out indirectly by application of a bandpass filter covering the range [0.5-3 Hz], then calculation of an energy of said resulting signal and subtraction from said calculated energy of said contribution of the area already expressed as an energy, which provides the energy of said residual signal corresponding to the noise level, here also expressed as an energy vector; - Calculation E3f of a signal-to-noise ratio by dividing the contribution by the noise level; a vector is then obtained with a value for each line corresponding to each area of interest; it should be noted that its values could be averaged (with or without weighting) at this or the next stage so that only one value to threshold is obtained at the next stage; - authorization (E6) to continue access to the data source based on the signal-to-noise ratio calculated here as a vector, a comparison is then made either to a single global threshold if only one value was obtained at the end of the signal-to-noise ratio calculation in the previous step, or by area of interest, the thresholds not being necessarily identical, which allows a sensitivity to be assigned by area (comparable to a weighting that would have been taken into account by area) and the sum of the binary values obtained per line, i.e. per area of interest, of the vector must exceed another threshold to allow continued access.
[0070] The illustrated embodiment is described in more detail in the following paragraphs.
[0071] The determination of the invariant parameters (not shown) by random selection of the value(s) of the invariant parameter(s) corresponds here to the random selection of the fundamental frequency of the imposed signal, identical for the two components blue and red, and of the phase shift between the blue and red components of the imposed signal. This results in the selection of a first value between 0.5 and 3 for the first invariant parameter (imposed fundamental frequency) and a random selection of a second value between 0 and 2ir for the second invariant parameter (phase shift between the two components, at least one parameter of which is time-invariant and whose value is randomly fixed: blue and red, since the green component, in the embodiment described here, is not time-variable but fixed). The random selection is implemented here by the remote server.
[0072] The draw described here is random; however, it can alternatively be pseudo-random, particularly so as not to repeat the same challenge on the same terminal and / or for the same person. To achieve this, the applied invariant parameters are stored in memory, preferably on the remote server side, in association with the identifier of each terminal (for example, the IMEI of a mobile phone) and / or the identifier (preferably anonymized) of each illuminated face for which the process has been applied, thus creating an exclusion register for subsequent draws. Preferably, this exclusion register corresponds to the external register of the time-stamped status, success (no fraud) or failure (fraud detected) of the security process. The terminal identifier (for example, the IMEI of a mobile phone) is, for example, communicated to the remote server as metadata during each connection to the remote server.The biometric identifier (preferably anonymized) of the illuminated face is, for example, created and stored, preferably by the remote server, by encrypting a biometric template obtained from the images acquired in the EL step. Thus, the next time an invariant parameter is drawn for the same terminal, the draw is pseudo-random because it excludes the values of parameters already applied to that terminal and / or for that user's face (by comparison, called matching, in a 1 vs. n ratio of the current biometric template with those already recorded in the exclusion register). Preferably, it is the combination of the same user on the same terminal that is excluded. This implementation method makes it possible, in particular, to control the case of a fraudster who might use different devices, potentially virtual ones.Furthermore, this exclusion register is advantageously monitored to track abnormal and potentially fraudulent implementations of the process. To this end, the invariant parameters applied are recorded in the register with their time-stamped status: success (no fraud) or failure (fraud detected). The number of failed attempts for the same biometric identifier over a given period triggers a system alert.
[0073] This imposed time-varying signal is a periodic signal composed of two periodic color components and constitutes a challenge, the periodic component(s) of the imposed time-varying signal being able to be sinusoidal, square wave or triangular, the two components not necessarily being of the same type.
[0074] If the periodicity of the imposed time-varying signal can be, as here, the same for all its components, it can also be due to only one component of the Red Green Blue display signal, namely the blue, green, or red component: the controlled time-varying signal is then composed of a single periodic color component, the others maintaining a fixed level during the control step, this level nevertheless being able to vary between each activation. The process involves these fixed levels being advantageously set randomly. However, it is preferable for the imposed time-varying signal to be composed of several color components, particularly periodic ones, to increase the diversity of the challenge, as is the case here, and the choice of only red and blue components (not green) helps to limit color crosstalk.
[0075] For example, the red and blue components of the imposed time-varying signal can be square wave and written as follows: = a mg (signelcoq 2^. / )) + 1) 'P' amètres ■ = «w„(sign(cos(2n- / Wœ q)+ 1) randomly are: - the red and blue frequencies in the working range [0.5Hz, 3Hz], The other parameters are fixed non-randomly; thus, the phase shift between the two components is set to 0. Similarly, the ared and abieu levels in the range are fixed non-randomly (no random selection) in the range [0, 255], preferably at 127.5. These amplitudes correspond to parameters that cannot be estimated in the observed signal (response of the captured object to the controlled lighting). Therefore, the amplitude levels are not randomly fixed invariant parameters.
[0076] Likewise, the green component has a fixed avert level (non-randomly) in the range [0, 255], preferably 255, but this level is not a randomly fixed invariant parameter, especially since the green component is not time-varying.
[0077] Preferably, the red and blue components of the imposed time-varying signal can be sinusoidal and written as follows: red = ared ( COS ( 2jrf rouget) + 1), WHERE the parameters are randomly set, blue = a bleu (cos() + 1) are : - the frequencies frouge and fHeu in the range [0.5Hz, 3Hz]; The other parameters are fixed non-randomly, so the phase shift between the two components is here fixed at 0. Furthermore, the ared and aMeu levels are fixed in the range [0, 255], preferably 127.5, as before, the green component then having an avert level fixed in the range [0, 255], preferably 255, as before.
[0078] In the mode illustrated in the embodiment specific to this figure, the red and blue components of the imposed time-varying signal are sinusoidal periodic and out of phase with each other and are written as follows: ( red = Ored ( cos ( 2ttfroug / ) + 1 , where the parameters are fixed \bleu = abieu(^2 +1) randomly are: - the frequency fred in the range [0.5Hz, 3 / ^We„Hz] so as to have the frequencies of the two components in the preferred interval, and - the phase in the range [ 0, 2ir ], which corresponds to the phase shift between the blue component and the red component of the imposed signal.
[0079] Moreover, the other parameters are fixed non-randomly: - the ared and abieu levels are fixed in the range [0, 255], preferably 127.5, as before, the green component then having an avert level fixed in the range [0, 255], preferably 255, as before; - and a natural number kbteu in the range [ 1, 6 ], is here fixed at 1. Thus, the frequencies of the two components are chosen to be equal (multiple = 1), which allows us to isolate a single frequency during step E3d of estimating the contribution, regardless of the component studied.
[0080] Fundamental frequencies above 3 Hz are preferentially avoided, both for user comfort and for sampling reasons, so that a video sequence can be processed as long as its sampling frequency is at least 6 Hz (Shannon criterion). However, it should be noted that mobile phones have cameras typically capable of recording video at at least 25 Hz locally (without external streaming) at high-density resolutions; therefore, the maximum value of this range can be increased, particularly if the acquired biometric information is not a face but a dermatoglyph, for example, as user comfort would then not be affected.
[0081] The E2 control of the display on a portion of the screen 5, with the color varying according to a time-varying sinusoidal signal of imposed fundamental frequency and phase shift between the blue and red components, the fundamental frequency and phase shift having been randomly fixed, allows the color in the portion of the screen 5 to vary periodically and without repeatability because it depends on two parameters drawn randomly and independently. Here, the red and blue components are periodic signals of fundamental frequency equal to the first randomly drawn value and having a phase shift between them equal to the second randomly drawn value, the green component being time-invariant (and initialized, for example, at level 255), the imposed signal thus encoding two invariant parameters which are here (non-exhaustively) parameters of the spectrum: the fundamental frequency of the imposed signal and the phase shift between the two time-varying components of the imposed signal. The display control E2 is performed here by user terminal 1 according to setpoint values to be applied per component sent in real time by the remote server, the time-varying red and blue sinusoids being calculated by the remote server based on the random draws it has performed, but not transmitted as such so as to keep the user terminal agnostic of the invariant parameters.
[0082] Advantageously, a preprocessing step (not illustrated) is applied here between the acquisition step E1 and the selection of the first series of images, notably by means of a preprocessing neural network. The main purpose of this step is normalization before face detection, to determine the area of interest, but it can also include color processing to abstract away luminance changes before the E3d estimation of the contribution. The neural network can be configured to address image compression artifacts and / or image sharpness. The neural network can also be configured to apply a colorimetric transformation to the original images. More generally, the image preprocessing performed can consist of one or more of the following image processing operators: - pixel-by-pixel (or point-by-point) modification operators. These include, for example, color, hue, and gamma correction; - local operators, especially those for managing local blur or contrast, a local operator based on a neighborhood of the pixel, that is to say more than a pixel but less than the whole image; a local operator allows, from a neighborhood of an input pixel, to obtain an output pixel; - operators in the frequency space (after image transformation). The use of one or more operators in the frequency space opens the way to various possibilities for analog or digital noise reduction, such as reducing compression artifacts, improving image sharpness, detail or contrast.
[0083] For the implementation of the evaluation step E3, and in particular the selection of images for the first series of images, a face detection step (not shown) is implemented here for all or part (after filtering out poor-quality images, for example) of the acquired images. Face detection includes determining a region of interest in an image corresponding to the location of the The face in the image. The region of interest is typically a rectangular area (or box) encompassing the face to isolate it from its surroundings in the image background. Other types of regions of interest can be used. For example, the region of interest can be defined by a boundary between the skin and the background, delineating the face from its background.
[0084] The region of interest can be detected using several approaches. One approach is, for example, to analyze the image to detect physical features such as the eyes. Since faces are made up of similar elements (eyes, noses, mouths, etc.) spatially organized in a similar way, the detection of these elements is facilitated. Another approach is, for example, to use a computational model such as a neural network, a support vector machine, or a decision tree, previously trained on a training set of images showing various faces.
[0085] This face detection step can be implemented by the information processing unit 106 on the images acquired and transmitted by the camera 3. It is also possible that this face detection step is implemented by another element than the information processing unit 106, such as the camera 3, and that the information processing unit 106 receives only the regions of interest rather than the complete images.
[0086] By way of exception, the optional alignment step E3a may consist of determining, using the same method as described in the preceding paragraph, a characteristic facial feature (eyes, nose, mouth, etc.) in each image and then aligning them with each other from one image to the next. Preferably, the eyes are chosen as alignment references because their alignment simultaneously allows for normalization and fixes the orientation of the coordinate system between the images, given that the neural network used for detecting the two eyes is preferentially trained to detect two eyes belonging to the same face, and in particular by discriminating between the right and left eyes so as to uniquely orient the images. Alternatively, and not by exception, a nose-mouth alignment is also possible but less accurate.
[0087] In the E3b step of determining at least two regions of interest, the regions of interest are preferentially geometrically defined from among: a triangle in a mesh of the detected face or a pixel in the vicinity of a detected facial landmark, each region of interest being present in all images of the first series. Such a landmark is, for example, a semantic point of the face, such as the tip of the nose, the corner of the eye (position sometimes estimated), as taught by Guo, X., Li, S., Yu, J., Zhang, J., Ma, J., Ma, L., & Ling, H. (2019). PFLD: A practical facial landmark detector. arXiv preprint arXiv: 1902.10859. Preferably, each region of interest This is an area of the face with a particularly distinctive shape and where there is the least possible distortion and / or obfuscation in order to reduce measurement noise during acquisition. Thus, the dorsum and the tip of the nose are particularly interesting areas, but not exclusively so.
[0088] The E3c step of determining a temporal signal observed per area of interest, representative of the pixel values of the images in the first series of images, is therefore based here only on a portion of the detected face, specifically on skin areas, using, for example, the segmentation method described in Wang, B., Chang, X., & Liu, C. (2011). Skin detection and segmentation of human face in color images, International Journal of Intelligent Engineering and Systems, 4(1), 10-17. This choice avoids biases related to eyeglasses, for example. For each image in the first series, a calculation of the average levels allows the creation of a temporal signal observed for said area of interest per component.
[0089] The E3d estimation step, by region of interest, of the contribution of the imposed emitted signal to the observed time-domain signal as a function of the two invariant parameters is performed component by component. The estimation of the contribution is implemented here for the fundamental frequency by bandpass time-domain filtering. To do this, and in a non-imitative manner, a narrow bandpass filter (to the tenth of a Hertz) centered on the imposed fundamental frequency (known fundamental frequency) is applied to the observed signal, and then the energy of the resulting filtered response is the contribution of the imposed emitted signal to the observed time-domain signal for the relevant component and region of interest. Choosing a bandpass filter with a shape slightly wider than the fundamental frequency of the imposed signal alone allows for the absorption of potential signal distortions by covering a slightly wider frequency spectrum.As an alternative to bandpass temporal filtering, a selection of the decomposition term according to the fast Fourier transform (FFT) of the acquired video temporal signal corresponding to the imposed (known) fundamental frequency is applied, which constitutes the estimation, already in the form of an energy, of the contribution of the imposed emitted signal to the observed temporal signal for the component and the region of interest concerned.
[0090] Another variant consists of correlating each component of the observed signal with a complex sinusoid of the imposed fundamental frequency (fred), that is, independently for each component, by putting the imposed signal into complex exponential form, which gives a vector r called imposed (with, for example: ) and a vector s called observed, then calculating the product The scalar of these two vectors is then divided by the magnitude of the imposed vector, and the resulting value constitutes the contribution c of the imposed emitted signal to the time-domain signal. observed for the component and the area of interest concerned, which can be written as:
[0091] Indeed, depending on the type of invariant, the method for estimating the contribution of the imposed emitted signal to the observed time signal, that is to say for isolating in the observed time signal the part which corresponds to the value of the known invariant parameter, differs: - for example in the case of a frequency invariant only, we advantageously isolate by frequency as described above, that is to say by means of frequency filtering, selection of an FFT term or by correlation by complex sinusoid; - for example, in the case of a phase-shift invariant between the components (for identical frequencies of the two components), the contribution is advantageously estimated by correlation with the phase-shifted complex sinusoids, using the same method as described previously by writing rk: = = for each component, then for example we sum these contributions, to finally calculate the modulus of this sum, which in the absence of fraud will be high (or alternatively we can for example normalize these vectors then subtract them and calculate the modulus of their difference, which in the absence of fraud will be almost zero or zero).
[0092] The E3e evaluation of a noise level of the time-domain video signal acquired by determination, by area of interest, is here carried out indirectly by applying a bandpass filter covering the working range [0.5-3 Hz], then calculating an energy of said resulting signal and subtracting from said calculated energy said contribution of the area already expressed as an energy, which provides the energy of said residual signal corresponding to the noise level, here also expressed as an energy vector with one row per area of interest and one column per blue, red component.
[0093] Step E3f of comparing the contribution with respect to the noise level here involves calculating a signal-to-noise ratio by dividing the contribution by the noise level.
[0094] In the implementation method using areas of interest, illustrated here, a calculation of the average, in particular weighted by area of interest, or of the median of said areas of interest so that the calculated signal-to-noise ratio is global, can notably be carried out as early as the contribution estimation step E3d and continued in the evaluation step E3e of a noise level of the acquired time-domain video signal; or during the calculation step E3f of a signal-to-noise ratio. In the implementation method described here, a weighted average calculation is preferably carried out in the calculation step E3f of the The signal-to-noise ratio is calculated from the noise level and contribution in matrix form. The optional use of a weighted average allows for different sensitivity levels to be applied to all areas of interest. This choice is particularly relevant when the areas cover a large portion of the region of interest, namely the face in the first image series. It also allows for a higher weighting to be assigned to an area of interest covering the imaged area of the eyes, particularly glasses, since glasses generate a strong return signal with low noise. This higher weighting can, for example, be applied only when glasses are detected in the first image series (for instance, using a neural network). Two weighting methods are described here, without limitation.
[0095] A first weighting method is based on an empirical qualification of the areas of interest according to criteria of stability in particular. For example, in the case of a face mesh composed of triangles, each triangle corresponding to an area of interest, each triangle is assigned a weighting value, and a lower value is assigned to the area(s) of interest overlapping what has been recognized in the image as a mouth, because a mouth is likely to be disturbed by movements.
[0096] Another weighting method is based on spatial contributions. This involves estimating normals to the surface of the acquired object (for example, according to the publication by Fu, Z., Hong, S., Liu, M., Laga, H., Bennamoun, M., Boussaid, F., & Guo, Y. (2023). Multi-stage information diffusion for joint depth and surface normal estimation. Pattern recognition, 141, 109660.) for each area of interest or for each pixel of the face.Then to calculate an expected spatial distribution (based on the postulate that the user's face is imaged frontally), for each area of interest or for each pixel of the face (map), as a function of the object-lighting distance and the orientation of the normal with respect to the object-lighting axis, which can be approximated by a constant distance and normal along the z-axis of the camera, and with a simplified lighting model (Lambertian) so that the expected spatial contribution = nz(x,y) with nz the z-component of the normal then we use this expected spatial contribution as a weighting.
[0097] Advantageously, a single overall signal-to-noise ratio value (averaged for all areas of interest and for both components) is then obtained by dividing: - the contribution in matrix form once averaged with the aforementioned weightings by - the noise level in matrix form once averaged with the aforementioned weightings.
[0098] Continued access to the data source is then permitted E6 if the previously obtained signal-to-noise ratio is greater than a given threshold, preferably a high one, which allows access to be permitted only in cases where the dominant frequency is that of the challenge, whereas a low threshold allows access even in cases where the challenge frequency is not dominant, as long as it is sufficiently represented. In practice, a threshold of 0.4 is chosen to operate with strong ambient lighting compared to the screen illumination 5.
[0099] Alternatively, in the case of frequency invariance, the E3 evaluation, based on a first series of images of the acquired time signal, of the presence of fraud by injection or presentation by: - E3d estimation of a contribution of the imposed emitted signal to the acquired video time signal as a function of said invariant parameter by determining the correlation rate for the imposed frequency (in particular by the component concerned) - evaluation E3e of a noise level of the acquired time-domain video signal by determining the correlation rate for each of the other frequencies; - comparison E3f of the contribution with respect to the noise level by determining the rank of the correlation rate for the imposed frequency.
[0100] By way of exception, this variant can be implemented as follows: for all or part of the operating frequency range, a frequency sampling interval is chosen (for example, 0.01 Hz), and for each frequency within the sampling interval, the similarity of a typical signal at that fundamental frequency to the received signal is determined. In other words, a discretization is performed: the acquired time-domain video signal is correlated, by periodic component, with the sinusoid of representative frequency (for example, the average frequency) of each sampling interval. This makes it possible to obtain a correlation coefficient (here in vector form) with each sampling frequency interval, that is, a similarity coefficient between the sinusoid of the acquired time-domain video signal and the sinusoid representing the sample. This step then corresponds to: - for the sampling interval including the frequency imposed on the E3d estimation of the contribution of the emitted signal imposed on the acquired video time signal as a function of the invariant parameters (which constitutes an interesting alternative to FFT, to avoid being dependent on the length of the signal, the number of images acquired and choosing the fineness of the step); - for the other sampling intervals at the E3e evaluation of the noise level of the acquired time-domain video signal, because all the correlation rates obtained except for that corresponding to the imposed (emitted) frequency correspond to the noise.
[0101] Then on the basis of these correlation rates, each sampling interval is classified or assigned an increasing rank going from the highest correlation rate to the lowest correlation rate, advantageously the rank is normalized by the number of intervals.
[0102] The E3f comparison of the contribution with respect to the noise level then consists of determining the rank proper to the sampling interval including the imposed frequency, for each component.
[0103] Finally, the E6 authorization to continue accessing the data source based on the comparison result consists, for example, of comparing the eigenrank of the sampling interval including the imposed frequency to a threshold rank. Thus, if the eigenrank is less than the threshold rank (for example, 0.03 in the case of a normalized rank), continued access to the data source is authorized. In the case of multiple time-varying imposed components, an overall rank can advantageously be determined by combining (for example, as a product or an average of the eigenranks of each component) the eigenranks determined in step E3f for each component, before thresholding the resulting overall rank, for example, at 0.001.
[0104] Alternatively, the method according to the invention could be applied to dermatoglyph capture, preferably using for acquisition the main camera of the terminal (rear of a "smartphone", or tablet for example): of better resolution and by making the flash blink according to the time-varying signal of imposed fundamental frequency and imposed phase.
[0105] With reference to [Fig. 4], the security process includes a supplementary verification by evaluation E4, based on a second series of images from the acquisition, for the presence of fraud by presentation by: - (optional) for each image of the second series, spatial alignment E4a of said images with respect to a reference point of the detected face, for example by eye registration, (or resumption of the results of the alignment step E3a if this step has been carried out); - determination E4b of at least two characteristic surfaces, in each image of the second series, each characteristic surface being common to the images of the second series, - determination E4c of an observed time-domain signal by characteristic surface, - estimation E4d, by characteristic surface, of a contribution of the imposed signal to the observed signal in order to form a feedback vector, - determination E4e of a score by applying a classifier to said vector, said classifier being implemented in particular by a neural network and the E6 authorization step of continuing access to the data source being a function of the determined score.
[0106] The E4 evaluation, based on a second set of images from the acquisition, for presentation fraud relies on verifying the object's properties by analyzing the acquired signal: that is, the received signal, which corresponds to the object's response (assumed to be the user's face) to the controlled lighting. This additional verification does not require estimating the challenge parameter in the acquired signal, nor any specific synchronization with a dedicated event to identify a dedicated sequence; therefore, there is no need to analyze the acquired images one by one, which makes the process flexible and memory-efficient. Finally, if the score indicates the presence of presentation fraud, access to the data source cannot continue.
[0107] The second series of images does not necessarily include all the images acquired in step El, as before, a selection of the images acquired to form said second series of images may be based for example on ISO quality criteria of the acquired images, or it may be possible to keep only the acquired images which include the entire face (as opposed to a face of which part is out of frame), the same criteria for selecting the images of the first series of images being able to be applied to the choice of the images of the second series of images, with regard to the characteristic surfaces.Similarly, the second series may consist only of images from a specific time window within the acquisition period E1 of camera 3. Preferably, this time window has a duration greater than or equal to twice the maximum period of the imposed signal(s) (per component), with a sampling frequency greater than or equal to twice the maximum frequency of the imposed signal(s) (per component). The second series of images is the same as the first series of images.
[0108] The characteristic surfaces belong to the region of interest imaging all or part of a face and correspond in particular to the surfaces on which the classifier was trained. These multiple characteristic surfaces are chosen so as to be able to highlight the differences in response related to the geometry and appearance of the object.They are not necessarily identical to the areas of interest, which may limit the reuse of the results of step E3b of determining the areas of interest; nevertheless, here the characteristic surfaces include the areas of interest of the dorsum and the tip of the nose as a high-response surface, and other characteristic surfaces corresponding to high-response surfaces are added: such as the imaged surface of the forehead or glasses and other characteristic surfaces corresponding to low-response surfaces are added: such as the imaged surface of the nostrils or the chin or the periauricular region.
[0109] The E4d estimation step, using a characteristic surface, of a contribution from the driven signal to the observed signal in order to form a feedback vector is implemented, for example, by filtering the observed signal at the imposed fundamental frequency and calculating the amplitude. In the presence of several characteristic surfaces, the amplitudes obtained per characteristic surface are then assembled together in the form of a feedback vector, also called an appearance vector. The feedback vector is then, for example, formed of as many sub-vectors as there are imposed time-varying signals, which allows the classifier to have a vector of vectors (also called a matrix, with, for example, as many columns as there are imposed signals (here, one column per periodic color component: blue and red) and as many rows as there are characteristic surfaces) as input.This feedback vector shows, for example, stronger responses at the level of the glasses or eyes, and variations depending on the parts of the face.
[0110] The invariant parameters (here the imposed fundamental frequency fred and the phase shift between the blue component and the red component) are known from the processing unit 106 implementing the estimation step E4d.
[0111] Then the step of determining E4e of a score involves the application of a classifier to this vector, said classifier being implemented in particular by a neural network.
[0112] Preferably, the classifier is said to be "deep" and constituted by supervised learning, previously trained from examples of real faces and frauds on screen or paper.
[0113] Then, the step of determining E4f the probability of fraud per presentation as a function of the determined score is implemented, for example by applying a sigmoid function or by using a table.
[0114] Finally, if the probability of fraud by presentation is less than a second predetermined threshold, this condition being preferentially concatenated with that described in step E6 with reference to the previous figure, i.e. that both conditions must be met to continue, then the continuation of the biometric authentication process for access to the account is authorized E6, and an accepted status of the request is recorded in a register linked to the security process, whereas otherwise the biometric authentication process is interrupted and a refused status of the request is recorded in the register linked to the security process.
[0115] This additional verification makes it possible to discriminate between screen fraud, i.e. if a fraudster positions a screen on which a video recording of the legitimate user is displayed, and paper fraud because then the return vector will not correspond to a correct return vector, for example because the responses will be substantially identical regardless of the areas, the screen or the paper presented, by the fraudster, facing the camera 3 being "flat".
[0116] With reference to [Fig. 5], in the case of a periodic imposed signal, the security process includes an additional check further improving robustness to specific recapture or injection, for example, by adding an evaluation step E5, from the third series of images of the acquisition, for the presence of specific fraud by recapture or injection by: -determination E5a of an observed latency as a function of the fundamental frequency of the imposed periodic time-variable signal and of an estimated phase shift between the imposed periodic time signal and an observed time signal representative of the pixel values of the images of the third series of images on all or part of the detected region of interest common to said images of the third series of images, - calculation E5b of a difference between the observed latency and a reference latency obtained previously for said terminal 1 or for the same terminal model, and the authorization E6 to continue access to the data source being a function of the calculated latency difference.
[0117] More specifically in this implementation mode: - The determination E5a of an observed latency is a function of the imposed fundamental frequency, here frmtge, and an estimated phase shift 0 between the imposed periodic time signal and the observed time signal, representative of the pixel values of the images in the third series of images on all or part of the detected face common to said images in the third series of images. Based on these two elements, a response time tréponse is then deduced to within one period: l , - ■—+ —2— with n: natural number, but equal to 0 by construction since, rÿpOflSC J rmt%e J Due to the imposed fundamental frequency range of the imposed signal, in practice the expected latency is on the order of 100ms, therefore significantly lower than Vf red, which allows the latency to be calculated precisely. - calculation E5b of a difference between the observed latency latobset a reference latency I at Ie[ obtained previously for said terminal or for the same terminal model. To obtain this reference latency, one can for example have a latency bank per terminal model, or consult the latency obtained by other users of the same device, or even have measured said reference latency during the enrollment of terminal 1, which is the simplest case but requires that terminal 1 has not changed in the meantime; - determination E5c of the specific probability of fraud by recapture or injection as a function of the calculated latency difference, for example using the function 1-d_^^£]with o a standard deviation (for example here of approximately 10ms), which allows farcr2 to assess the likelihood of the observed latency if we assume that the latency observed has a Gaussian distribution (normal law) centered on the reference latency with a standard deviation of 0.
[0118] Finally, if the specific probability of fraud by recapture or injection is less than a third predetermined threshold, this condition being concatenated to that described in step E6 with reference to the previous figures, then the continuation of the biometric authentication process for access to the account is authorized E6, and an accepted status of the request is recorded in a register linked to the security process, whereas otherwise the biometric authentication process is interrupted and a refused status of the request is recorded in the register linked to the security process.
[0119] Alternatively, as with steps E4 and E3, the intermediate probability calculation step is not limiting, the access authorization of step E6 may depend on the inferiority of the difference in latencies calculated to a threshold of approximately 10ms for example.
[0120] The image series used here is the first image series.
[0121] This additional step makes it possible to verify the consistency between the observed latency and the reference latency and thus to better identify expert fraud by recapture or injection.
[0122] Preferably, an acquisition parameter, such as the camera image size, is modified during the acquisition El and the specific probability of fraud by recapture or injection is then increased in the absence of a change in the latency observed during the acquisition, because this means that the challenge is not checked since the change in image size necessarily changes the capture delay.
[0123] By way of non-limitation, the first, second and third thresholds are for example set at 0.5.
[0124] By way of non-limitation, in figures 3 to 5 the video acquisition step E1 by camera 3 has been positioned upstream of the control step E2, but the two can be concomitant, or even the control step E2 could launch the video acquisition E1 by camera 3.
Claims
Demands
1. A method (P) for securing remote access to a data source from a terminal (1) by a user of said terminal, the terminal comprising a light source (5) and a camera (3) directed towards said user, said method comprising the steps of: - acquisition (E1) of a time-domain video signal by the camera (3); - control (E2), during the acquisition step, of the light source (5) according to a time-varying, periodic, imposed signal, at least one parameter of which is time-invariant and of randomly fixed value; - evaluation (E3), from a first series of images of the acquired time-domain signal, of the presence of fraud by injection or presentation by: - estimation (E3d) of a contribution of the imposed emitted signal to the acquired time-domain video signal as a function of said parameter; - evaluation (E3e) of a noise level of the acquired time-domain video signal; - comparison (E3f) of the contribution with respect to the noise level;- authorization (E6) to continue access to the data source depending on the result of the comparison step.;
2. A method according to the preceding claim, wherein the imposed time-variable emitted signal is a periodic color component whose time-invariant parameter is a fundamental frequency.
3. A method according to any one of the preceding claims, wherein the imposed time-variable emitted signal comprises at least two periodic color components of the same fundamental frequency or multiples of the same fundamental frequency, with a phase shift between said at least two components, said at least one time-invariant parameter of the imposed signal being said fundamental frequency or the phase shift.
4. A method according to any one of the preceding claims, wherein the imposed fundamental frequency is randomly fixed within a working range between 0.5Hz and 3Hz.
5. A method according to any one of the preceding claims, wherein the terminal comprises a display screen, the source luminous being all or part of the display screen (5) of the terminal (1) and the control of the light source corresponding to a display on all or part of the screen (5) of color varying according to the time-varying signal imposed.
6. A method according to any one of the preceding claims, wherein the estimation (E3d) of the contribution includes a bandpass temporal filtering or a predetermined component selection of a fast Fourier transform of the acquired video temporal signal, or a correlation with the imposed emitted signal.
7. A method according to any one of the preceding claims, wherein the contribution is expressed in the form of an energy, the evaluation (E3e) of the noise level of the acquired time-domain video signal being implemented by calculating an energy of a residual signal corresponding to a subtraction of a time average of the acquired time-domain video signal from said acquired time-domain video signal.
8. A method according to any one of the preceding claims, wherein the step (E3f) of comparing the contribution with respect to the noise level comprises a calculation of a signal-to-noise ratio by dividing the contribution by the noise level.
9. A method according to any one of the preceding claims, wherein at the evaluation step (E3), from a first series of images of the acquired time signal, of the presence of fraud by injection or presentation: - at least one area of interest is determined (E3b), in each image of the first series, said area of interest being common to the images of the first series, - an observed time signal representative of the pixel values of the images of the first series of images over all or part of the area of interest is determined (E3c); - during the estimation (E3d) said time signal of the first series of images is said observed time signal determined by area of interest.
10. Method according to the preceding claim, wherein the step of - evaluation (E3e) of a noise level of the acquired time-domain video signal; or - comparison (E3f) of the contribution with respect to the noise level;
11.
12.
13. includes a calculation of the average, in particular weighted by area of interest, or of the median of said areas of interest so that the calculated signal-to-noise ratio is global. A method according to any one of the preceding claims, comprising an evaluation step (E4), based on a second series of images from the acquisition, for the presence of fraud by presentation by: - determination (E4b) of at least one characteristic surface, in each image of the second series, said characteristic surface being common to the images of the second series, - determination (E4c) of a temporal signal observed by a characteristic surface, - estimation (E4d), by characteristic surface, of a contribution of the imposed signal to the observed signal in order to form a feedback vector, - determination (E4e) of a score by applying a classifier to said vector, said classifier being implemented in particular by a neural network and the authorization step (E6) to continue access to the data source being a function of the determined score. A method according to any one of claims 9 to 11, wherein the area of interest or the characteristic surface belongs to an area of interest imaging all or part of a face. A method according to any one of the preceding claims, comprising an evaluation step (E5), from a third series of images from the acquisition, of the presence of specific fraud by recapture or injection by: -determination (E5a) of an observed latency as a function of the fundamental frequency of the imposed periodic time-varying signal and of an estimated phase shift between the imposed periodic time signal and an observed time signal representative of the pixel values of the images of the third series of images over all or part of the detected region of interest common to said images of the third series of images, - calculation (E5b) of a difference between the observed latency and a reference latency previously obtained for said terminal (1) or for the same terminal model, and the authorization (E6) to continue access to the data source being a function of the calculated latency difference.
14. A computer program comprising instructions adapted to the implementation of each of the steps of the method for securing remote access to a data source according to any one of claims 1 to 13 when said program is executed on a computer.
15. Non-transient information storage means, removable or not, partially or totally readable by a computer or microprocessor, comprising code instructions of a computer program for the execution of each of the steps of the process according to any one of claims 1 to 13.
Citation Information
Patent Citations
Method for analyzing a facial characteristic of a face
FR3100074A1
System and method for authorizing access to access-controlled environments
US20150195288A1
Systems and methods for liveness analysis
US20160071275A1
System and method for extracting a periodic signal from video
US20180122066A1
FR1661737S