DEVICE AND METHOD FOR RETINAL AUTHENTICATION AND IDENTIFICATION
A refined model for retinal authentication and identification is developed using sister and cousin images through self-supervised and supervised training, addressing the challenges of feature extraction and acquisition variations, providing robust and secure authentication.
Patent Information
- Application Number
- FR2024004149
- Authority / Receiving Office
- FR · FR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-22
- Publication Date
- 2025-10-24
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Current retinal image authentication methods face challenges due to the complexity of extracting and comparing image characteristics, such as the vascular network, which are not robust enough for reliable identification and can be easily copied, and are affected by variations in retinal images due to different acquisition conditions.
A refined model is developed using a self-supervised and supervised training approach that generates characteristic vectors directly from retinal images without extracting specific features, by creating sister and cousin images through various image processing techniques to enhance data availability and robustness, utilizing a Siamese network for authentication.
The refined model provides robust authentication and identification by being insensitive to variations in image acquisition, reducing the need for feature extraction and comparison, thus enhancing security and accuracy.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Title of the invention: DEVICE AND METHOD FOR RETINAL AUTHENTICATION AND IDENTIFICATION FIELD OF THE INVENTION
[0001] The present invention relates to the field of secure authentication and identification of an individual.
[0002] More specifically, the present invention relates to computer-implemented devices and methods for enabling the authentication and / or identification of an individual from an image of their retina and a refined model previously obtained by training a pre-trained foundation model on a training database constructed from a set of raw retinal images acquired from a plurality of subjects. STATE OF THE ART
[0003] Reliable authentication of individuals is a major concern in the field of IT security. Traditional methods, such as the use of passwords, have vulnerabilities related to password management and brute force attacks. Therefore, there is a growing need for innovative and secure solutions to robustly authenticate users.
[0004] In recent years, devices and methods for identifying and authenticating individuals have evolved towards the use of biometric markers such as fingerprints, face or even the iris of the eye. However, these markers are still easily accessible and a malicious person could copy them to steal the identity of another person.
[0005] The use of an individual's retinal image, which requires a specific imaging system, makes it a particularly interesting biometric marker since this retinal image cannot be acquired without the explicit consent of the individual to be authenticated. Current authentication methods using this type of image are based on the extraction and comparison of image characteristics such as the vascular network for example. However, the extraction of characteristics from retinal images for authentication or identification purposes is complex and does not currently provide sufficient performance for the desired purpose. Furthermore, when comparing two retinal images, two technical problems must be taken into account: two images of the same retina do not necessarily represent the same area; and - the retinal imprint may have changed between two shots.
[0006] The present invention proposes to match retinal images capable of overcoming the obstacles identified above. In particular, the invention makes it possible to avoid the step of extracting and prior recognition of image characteristics. SUMMARY
[0007] A first aspect of the invention relates to a device for obtaining a refined model (called “fine-tuned model” in English) configured to allow the authentication of an individual from at least one retinal image acquired on an eye of said individual.
[0008] This device comprises:
[0009] at least one input configured to receive: - a foundation model configured to, from a received retinal image, generate at least one characteristic vector associated with the received retinal image; - a first set of raw retinal images acquired from a plurality of subjects and comprising one raw retinal image per subject; - a second set of raw retinal images acquired from a plurality of subjects, comprising at least two raw retinal images of the same eye for each subject;
[0010] at least one processor configured to: - for each of the raw retinal images of said first set of images, generating at least one sister retinal image by applying at least one image processing to said raw retinal image and associating with said sister image thus generated the identity of the subject from which said associated raw retinal image originates; - generating a first training database comprising the raw retinal images of said first set and their generated sister retinal images; - pre-train the foundation model with the first training database, in a self-supervised manner, by defining similarity relationships between sister retinal images; - for each of the raw retinal images of said second set of images, generating at least one sister retinal image by applying at least one image processing to said raw retinal image and associating with said sister image thus generated the identity of the subject from which said associated raw retinal image originates; - generating a second training database comprising the raw retinal images of said second set and their generated sister retinal images, where any pair of non-sister retinal images sharing an identity of the same subject are said to be cousins; - obtaining said refined model by training said pre-trained foundation model on the second training database, in a supervised manner, by defining similarity relationships between the cousin retinal images;
[0011] at least one output configured to provide said refined model.
[0012] The device thus obtained makes it possible to obtain a refined model which, thanks to the training of artificial intelligence, for each retinal image received, generates at least one associated characteristic vector without having to go through extraction steps, such as segmentation, or recognition of particular characteristics present in the captured retina of the image, such as the vascular network. The generation of sister and cousin images also makes it possible to artificially increase the number of images to pre-train the foundation model and then refine it.
[0013] This global strategy, making it possible to obtain the refined model robust to the different acquisition conditions modeled by the augmentations, has multiple advantages, such as: the availability of a large set of retinal images (i.e., the first set of raw retinal images which includes a single raw retinal image per subject); in fact, most public databases of raw retinal images only include a single image per identity, i.e., associated with the same retina. The present invention presents a strategy for creating training databases (i.e., first training database and second training database) and global training (pre-training plus refinement) which makes it possible to overcome this initial lack of data.
[0014] The pre-training of the foundation model is therefore done on a first database of sister images (i.e., a database obtained from the first set of raw retinal images and comprising subsets of at least two sister images, one real and the other artificial, associated with the same identity), so as to make it possible to keep the second set of raw retinal images for the model refinement phase. In addition, the pre-training which is carried out in a self-supervised manner, has the advantage of making it possible to obtain a pre-trained foundation model having acquired general knowledge on the structure and texture of the retinal images.
[0015] The present invention then proposes the refinement (called “fine tuning” in English) of the pre-trained foundation model on the second training database, also augmented, but more restricted because it is built from from the second set of raw retinal images, comprising at least two raw retinal images of the same identity. Indeed, few such data (i.e., multiple retinal images of the same identity) are currently accessible, therefore a strategy allowing the model to be refined on a restricted training database is highly advantageous.
[0016] Preferably, pre-training the foundation model with the first training database is performed by also defining dissimilarity relationships between the non-sister retinal images, possibly allowing the use of a contrastive cost function based on a metric (such as Euclidean distance).
[0017] Preferably, obtaining the refined model by training the pre-trained foundation model on the second training database, in a supervised manner, is achieved by also defining dissimilarity relationships between the non-cousin and non-sister retinal images.
[0018] Preferably, obtaining the refined model by training, in a supervised manner, the pre-trained foundation model on the second training database is also done so as to minimize a contrastive cost function from the similarity relationships defined between the cousin retinal images and the dissimilarity relationships defined between the non-cousin retinal images.
[0019] Preferably, the at least one processor is also configured to:
[0020] - receive a validation base comprising a third set of images raw retinal images of a group of individuals, each individual being identified by at least two raw retinal images of the same identity of an individual;
[0021] - validate said refined model using the validation base.
[0022] Preferably, the at least one processor is also configured to apply at least one image pre-processing to the received retinal images, among which: a change of color space, a resizing, a cropping, a filtering. These different image pre-processings allow, among other things: the use of grayscale images, to clean the image of certain artifacts, to recover only the retinal disc (useful part of the image) and to ensure that all the retinal images are at the same working resolution.
[0023] Preferably, the at least one image processing applied to generate the sister retinal images comprises a color change, said color change comprising at least one of: a histogram matching between the retinal image and a reference retinal image randomly chosen from a set of reference retinal images, a random modification of the channels of the retinal image in a color space (for example: the “Red, Green, Blue” or RGB space, the “Cyan, Magenta, Yellow” or CMY space which can be completed by an additional black or K component, the "Hue, Saturation, Value" or HSV space, or the "Luminance, a, b" or LAB space), a conversion of the retinal image into black and white. Since images can be acquired by different imaging systems, this type of image processing makes it possible to make the models insensitive (i.e. more robust) to variations linked to image acquisition, in order to improve the performance of the foundation model and the refined model.
[0024] Preferably, the at least one image processing applied to generate the sister retinal images comprises an elastic deformation of the retinal image by applying a predefined displacement field randomly chosen from a set of predefined displacement fields, preferably of the diffeomorphic type. This makes it possible to simulate small changes in the structure of the vascular network that may occur due to different hardware parameters or acquisition fields.
[0025] Preferably, the at least one image processing applied to generate the sister retinal images comprises a change in brightness comprising at least one of: an overall darkening of the retinal image, a local overexposure of the retinal image, an addition of a darkening spot. This makes it possible to simulate the effect on the fundus images of variations in illumination induced by poor lighting environments and equipment, moving individuals or insufficiently dilated pupils.
[0026] Preferably, the at least one image processing applied to generate the sister retinal images comprises a degradation of the retinal image by the addition or removal of at least one of: noise, blurring, desaturation. This makes it possible to bring together the transformations degrading the quality of the acquisition such as noise, blurring or desaturation.
[0027] Preferably, at least one of the sister retinal images generated from the raw retinal images of the first set of images is generated by applying at least one translation of the retinal image. This transformation, although unrealistic, makes it possible to break the obvious biases of the images such as the cropping or the relative position of the optic disc within the retinal fundus disc.
[0028] Preferably, the foundation model comprises at least one of: a transformer (called "Transformer" in English), a convolutional neural network. This makes it possible in particular to use current SOTA computer vision techniques.
[0029] Preferably, said refined model comprises at least one Siamese network composed of at least two identical convolutional neural networks obtained by training at least two identical foundation models pre-trained on the first training database, then trained on the second training database.
[0030] Preferably, said at least one characteristic vector generated for each retinal image received has a dimension greater than or equal to 5. This makes it possible to encode a minimum of information.
[0031] Preferably, at least two sister retinal images are generated for each raw retinal image received from the first set of images, preferably at least ten. Since the augmentation process can be quite slow and cannot be done on the fly during training, this makes it possible to have a minimum of sister images previously generated statically (for example: stored on the hard disk).
[0032] A second aspect of the invention relates to an authentication device for enabling the authentication of an individual from at least one retinal image acquired on at least one eye of the individual and from a refined model obtained using the device previously described.
[0033] The authentication device comprises:
[0034] at least one input configured to receive: - at least one retinal image to be authenticated acquired on the individual; - the refined model configured for, from a received retinal image, generating at least one feature vector associated with the received retinal image;
[0035] at least one processor configured to: - generate, using the refined model, at least one vector of characteristics from the at least one retinal image to be authenticated; - calculate at least one estimated similarity score between at least one reference vector associated with the individual and at least one characteristic vector relating to the retinal image to be authenticated; - authenticate the at least one retinal image to be authenticated as belonging to the individual from an authentication condition based on the at least one similarity score;
[0036] at least one output configured to provide as output at least one of: the at least one calculated similarity score, a result relating to the authentication of the at least one retinal image to be authenticated, a command relating to an action to be performed.
[0037] The authentication device thus obtained makes it possible to authenticate an individual without having to go through steps of extracting characteristics, such as segmentation, and / or recognizing and then comparing particular characteristics of the image such as the vascular network. Indeed, only the vector of characteristics associated with the image to be authenticated is compared to a reference vector.
[0038] Preferably, the at least one input is also configured to receive at least one previously acquired reference retinal image of the individual, and the at least one processor is further configured to generate, using the refined model, the at least one reference vector associated with the individual from the at least one reference retinal image.
[0039] Preferably, the at least one processor is also configured to apply at least one image pre-processing to the received retinal images, among which: a change of color space, a resizing, a cropping, a filtering. As stated previously, these different image pre-processings allow, among other things: the use of grayscale images, to clean the image of certain artifacts, to recover only the retinal disc (useful part of the image) and to ensure that all the retinal images are at the same working resolution.
[0040] Preferably, the similarity score is estimated from the contrastive cost function of the refined model.
[0041] Preferably, the at least one similarity score is calculated from at least one distance calculated between at least one reference vector associated with the individual and the at least one authentication vector relating to the retinal image to be authenticated, preferably the at least one distance is a Euclidean distance.
[0042] Preferably, the authentication condition comprises the comparison of the at least one similarity score with at least one predefined threshold, said at least one predefined threshold being calculated following the application of a cross-validation method carried out on the refined model using a validation base comprising a third set of raw retinal images of a group of individuals, each individual being identified by at least two raw retinal images of the same identity.
[0043] A third aspect of the invention relates to a device for identifying an individual from at least one retinal image acquired on at least one eye of said individual, from a refined model obtained using the device previously described.
[0044] The identification device comprises:
[0045] at least one input configured to receive: - at least one retinal image acquired on said individual to be identified; - said refined model configured to, from a received retinal image, generate at least one vector of characteristics associated with the received retinal image;
[0046] at least one processor configured to: - generating, using said refined model, at least one characteristic vector from said at least one received retinal image; - calculate at least one estimated similarity score between a set of reference vectors associated with a set of individuals whose identity is known and said at least one vector of characteristics relating to the received retinal image; - determining the identity of said individual from among the set of individuals whose identity is known, from a condition based on said at least one similarity score;
[0047] at least one output configured to provide as output at least one of: said at least one calculated similarity score, a result relating to the identity of said individual associated with said at least one received retinal image, a command relating to an action to be carried out.
[0048] A fourth aspect of the invention relates to a computer-implemented method for obtaining a refined model configured to enable the authentication of an individual from at least one retinal image acquired on an eye of said individual.
[0049] The method of obtaining a refined model comprises: - receiving a foundation model configured to, from a received retinal image, generate at least one feature vector associated with the received retinal image; - receiving a first set of raw retinal images acquired from a plurality of subjects and comprising one raw retinal image per subject; - receiving a second set of raw retinal images acquired from a plurality of subjects, comprising at least two raw retinal images of the same eye for each subject; - for each of the raw retinal images of said first set of images, generating at least one sister retinal image by applying at least one image processing to the raw retinal image and associating with the sister image thus generated the identity of the subject from whom the associated raw retinal image originates; - generating a first training database comprising the raw retinal images of said first set and their generated sister retinal images; - pre-train the foundation model with the first training database in a self-supervised manner by defining similarity relationships between sister retinal images; - for each of the raw retinal images of the second set of images, generating at least one sister retinal image by applying at least one image processing to the raw retinal image and associating with the image sister thus generated the identity of the subject from which the associated raw retinal image comes; - generating a second training database comprising the raw retinal images of the second set and their generated sister retinal images, where any pair of non-sister retinal images sharing an identity of the same subject are said to be cousins; - obtain the refined model by training the pre-trained foundation model on the second training database, in a supervised manner, by defining similarity relationships between the cousin retinal images; - provide the refined model.
[0050] The preferential and advantageous characteristics related to the device previously described are also applicable to the method for obtaining a refined model described below.
[0051] A fifth aspect of the invention relates to a method for authenticating an individual from at least one retinal image acquired on at least one eye of said individual and from a refined model obtained according to the method described above.
[0052] The authentication method can be implemented by a computer and comprises: - receiving at least one retinal image to be authenticated acquired on said individual; - receiving said refined model configured to, from a received retinal image, generate at least one characteristic vector associated with the received retinal image; - generating, using said refined model, at least one vector of characteristics from said at least one retinal image to be authenticated; - calculate at least one estimated similarity score between at least one reference vector associated with said individual and said at least one characteristic vector relating to the retinal image to be authenticated; - authenticate said at least one retinal image to be authenticated as belonging to said individual from an authentication condition based on said at least one similarity score; - provide at least one of: said at least one calculated similarity score, a result relating to the authentication of said at least one retinal image to be authenticated, a command relating to an action to be carried out.
[0053] The preferential and advantageous characteristics related to the method for obtaining a refined model previously described are also applicable to the authentication method described below.
[0054] A sixth aspect of the invention relates to a method implemented to identify an individual from at least one retinal image acquired on at least one eye of said individual and from a refined model obtained according to the method previously described.
[0055] The identification method comprises: - receive at least one retinal image acquired on the individual to be identified; - receive the refined model configured to, from a received retinal image, generate at least one characteristic vector associated with the received retinal image; - generate, using the refined model, at least one vector of characteristics from the at least one retinal image to be authenticated; - calculate at least one estimated similarity score between a set of reference vectors associated with a set of individuals whose identity is known and the at least one vector of characteristics relating to the received retinal image; - determining the identity of the individual among the set of individuals whose identity is known, from a condition based on at least one similarity score; - provide at least one of: at least one calculated similarity score, a result relating to the identity of the individual associated with the at least one retinal image received, a command relating to an action to be carried out.
[0056] The preferential and advantageous characteristics related to the method for obtaining a refined model previously described are also applicable to the identification method described below.
[0057] A seventh aspect of the invention relates to a computer program product comprising instructions which, when the program is executed by a computer, causes the computer to perform any of the methods previously described.
[0058] An eighth aspect of the invention relates to a non-transitory computer-readable storage medium comprising instructions which, when the medium is read by a computer, cause the computer to perform any of the methods previously described. DEFINITIONS
[0059] In the present invention, the terms below are defined as follows: - The terms "subject" and "individual" refer to a living organism, such as a human being, an animal or an experimental subject, on which the invention is intended to be used or tested. In the present case, it is preferably a human. - The terms “adapted” and “configured” are used in this disclosure to broadly encompass the initial configuration, the further adaptation or supplementation of the present device, or any similar combination thereof, whether by hardware or software (including firmware) means. The term "processor" should not be construed as being limited to hardware capable of executing software, and generally refers to a processing device, which may, for example, include a computer, a microprocessor, an integrated circuit, or a programmable logic device (PLD). The processor may also include one or more graphics processing units (GPUs), whether for computer graphics and image processing or other functions. In addition, the instructions and / or data for performing the associated and / or resulting functionality may be stored on any processor-readable medium such as an integrated circuit, a hard disk drive, a CD (Compact Disc), an optical disk such as a DVD (Digital Versatile Disc), RAM (Random-Access Memory), or ROM (Read-Only Memory).Instructions may be stored in hardware, software, firmware, or any combination thereof. Machine learning (ML) traditionally refers to computer algorithms that automatically improve through experience, based on training data, to adjust the parameters of computer models by reducing the gaps between the expected outputs extracted from the training data and the evaluated outputs calculated by the computer models. In ML applications, a "hyperparameter" is a parameter of a model that is not optimized during training but is used to control the learning process, such as the mini-batch size and learning rate of a gradient descent algorithm, the width of a time window, or the number of convolutional layers of a convolutional network. "Datasets" are collections of data used to build an ML mathematical model, in order to make predictions or decisions based on the data. In "supervised learning" (i.e., the inference of functions from known input-output examples in the form of labeled or annotated training data), three types of ML datasets (also referred to as ML ensembles) are generally dedicated to three respective types of tasks: "training", i.e., parameter adjustment, "validation", i.e., adjustment of ML hyperparameters (which are parameters used to control the learning process), and "testing", i.e., independent verification of a training data set used to build a mathematical model that the latter provides satisfactory results. - A "neural network" (NN) refers to a category of ML comprising nodes (called "neurons") and connections between neurons modeled by "weights". For each neuron, an output is given as a function of an input or set of inputs by an "activation function". Neurons are generally organized into several "layers", so that neurons in one layer only connect to neurons in the immediately preceding and immediately following layers.
[0060] The above ML definitions are consistent with their usual meaning, and may be supplemented by numerous associated features and properties, as well as related digital object definitions, well known to a person skilled in the ML field. BRIEF DESCRIPTION OF THE FIGURES
[0061] The present disclosure will be better understood, and other specific features and advantages will become apparent, upon reading the following description of particular and non-restrictive illustrative embodiments, the description making reference to the accompanying drawings where:
[0062] [Fig.l] is a flowchart illustrating the different steps of a computer-implemented method for obtaining a refined model according to one embodiment of the invention.
[0063] [Fig.2] represents a possible embodiment for a device which can be used to obtain a refined model, authenticate and / or identify an individual.
[0064] [Fig. 3] represents a first set of raw retinal images acquired on a plurality of individuals, as well as a first training database generated from this first set and comprising sister images from the same raw image.
[0065] [Fig.4] represents a second set of raw retinal images acquired on a plurality of individuals, as well as a second training database generated from this second set and including sister images from the same raw image and cousin images from at least two different raw images of the same identity (i.e. the same eye) of an individual.
[0066] [Fig.5] represents a flowchart illustrating the different image processing operations applicable to generate sister retinal images from raw retinal images according to one embodiment of the invention.
[0067] [Fig.6] represents retinal images having a great colorimetric diversity.
[0068] [Fig.7] represents an example of histogram matching applied to a raw retinal image (left image) from a reference retinal image (bottom image) taken randomly from the set of retinal images of [Fig.6], to obtain a sister retinal image (top right image).
[0069] [Fig.8] shows an example of stationary velocity fields applicable as elastic deformation to generate sister retinal images.
[0070] [Fig.9] shows examples of sister retinal images obtained by applying elastic deformation to the original retinal image (located at the top left), such as those illustrated in [Fig.8].
[0071] [Fig. 10] shows examples of retinal images including variations in illumination induced by environments or equipment with poor lighting, individuals moving during acquisition of the retinal image or pupils that are not sufficiently dilated.
[0072] [Fig.11a] and [Fig. 11b] represent two examples of illumination maps modeling illumination variations applicable to simulate brightness changes and which can be used to generate sister retinal images.
[0073] [Fig. 12] represents examples of sister images obtained after applying a change in brightness to an original retinal image (located at the top left).
[0074] [Fig. 13] shows an example of degradations, such as adding / removing blur, applicable to an original retinal image (located at the top).
[0075] [Fig. 14] represents an example of translation applicable to an original retinal image, in particular to the optic disc (also called optic nerve) present on the original retinal image, to obtain a sister retinal image.
[0076] [Fig. 15] illustrates the principle of forming triplets which are used for the triplet contrastive cost function (called “triplet loss function” in English).
[0077] [Fig. 16a], [Fig. 16b], [Fig. 16c] and [Fig.16d] illustrate an example of image preprocessing applicable to raw retinal images.
[0078] [Fig. 17] represents a flowchart illustrating the different steps of a computer-implemented method for authenticating an individual according to one embodiment of the invention.
[0079] [Fig. 18] represents a flowchart illustrating the different steps of a computer-implemented method for identifying an individual according to one embodiment of the invention.
[0080] In the figures, the drawings are not to scale and identical or similar elements are designated by the same references. DETAILED DESCRIPTION
[0081] The present description illustrates the principles of the present disclosure. It is therefore obvious that those skilled in the art will be able to devise various arrangements which, although not explicitly described or illustrated herein, embody the principles of the disclosure and are included within its spirit and scope.
[0082] All examples and conditional terms cited herein are intended to assist the reader in understanding the principles of the disclosure and the concepts contributed by the inventor to the advancement of the art, and should be construed as not being limited to the specifically cited examples and conditions.
[0083] Furthermore, all statements of principles, aspects, and embodiments of the disclosure, as well as specific examples, are intended to encompass structural and functional equivalents thereof. Furthermore, such equivalents are intended to include both currently known equivalents and equivalents developed in the future, i.e., any developed element that performs the same function, regardless of structure.
[0084] Thus, for example, those skilled in the art will appreciate that the block diagrams presented herein may represent conceptual views of illustrative circuits incorporating the principles of the disclosure. Similarly, it will be appreciated that all flowcharts, flow diagrams, and the like represent various processes that may be substantially represented on a computer-readable medium and thereby executed by a computer or processor, whether or not that computer or processor is explicitly shown.
[0085] The functions of the various elements illustrated in the figures may be provided by the use of dedicated hardware or hardware capable of executing software in association with appropriate software. When provided by a processor, the functions may be provided by a single dedicated processor, a single shared processor, or several individual processors, some of which may be shared.
[0086] It is to be understood that the elements illustrated in the figures may be implemented in various forms of hardware, software, or combinations thereof. Preferably, these elements are implemented in a combination of hardware and software on one or more suitably programmed general-purpose devices, which may include a processor, memory, and input / output interfaces.
[0087] DEVICE AND METHOD FOR OBTAINING AN AFFINITE MODEL
[0088] According to a first aspect, the present disclosure relates to a device and a method implemented by computer for obtaining a refined model (called "fine-tuned model" in English) configured to allow the authentication of an individual from at least one retinal image acquired on an eye of the individual (each individual having an identity per eye).
[0089] As illustrated in [Fig.l], the method of obtaining a refined model may comprise the following main steps (which will be described in more detail later): - step S10: receiving a foundation model configured to, from a received retinal image, generate at least one characteristic vector associated with the received retinal image; - step S20: receiving a first set of raw retinal images El acquired from a plurality of subjects and comprising one raw retinal image per subject; - step S30: receiving a second set of raw retinal images E2 acquired from a plurality of subjects, comprising at least two raw retinal images of the same eye for each subject; - step S40: for each of the raw retinal images of the first set of images E1, generate at least one sister retinal image by applying at least one image processing to the raw retinal image and associate with the sister image thus generated the identity of the subject from which the associated raw retinal image originates; - step S50: generating a first training database B1 comprising the raw retinal images of the first set El and their generated sister retinal images; - step S60: pre-train the foundation model with the first training database B1, in a self-supervised manner, by defining similarity relationships between the sister retinal images; - step S70: for each of the raw retinal images of the second set of images E2, generating at least one sister retinal image by applying at least one image processing to the raw retinal image and associate with the sister image thus generated the identity of the subject from which the associated raw retinal image originates; - step S80: generating a second training database B2 comprising the raw retinal images of the second set E2 and their generated sister retinal images, where any pair of non-sister retinal images sharing an identity of the same subject are said to be cousins; - step S90: obtaining said refined model by training said pre-trained foundation model on the second training database B2, in a supervised manner, by defining similarity relationships between the cousin retinal images; - step S120: provide said refined model.
[0090] In step S10 of receiving the foundation model, the foundation model (called "foundation model" in English) is an artificial intelligence model that is used as a starting point for the fine-tuning process. Said foundation model can be large, and therefore require a large amount of data to be pre-trained. Advantageously, the received foundation model may have been previously initialized (i.e., pre-trained before the pre-training defined in step S60) using a database comprising a set of any images (i.e., not necessarily retina images) such as, for example, the ImageNet database which is often used in the field of deep learning and which contains a vast collection of real images. Foundation models that already exist and are "ready to use" (i.e., have already been initialized), such as those available in the Keras library, can also be used.
[0091] Advantageously, the foundation model may comprise at least one of: - a transformer (called “Transformer” in English) such as a ViT (for “Vision Transformer” in English) whose architecture is particularly suited to image classification, or a Swin (for “Swin Transformer” in English) which has notably demonstrated competitive performance on several computer vision tasks including image classification on ImageNet; - a convolutional neural network (called “Convolutional Neural Network” or CNN in English) such as: • a ResNet (for “Residual Network” in English) which introduces the concept of “shortcuts” or residual connections which allow the direct passage of information from one layer to another, facilitating the training of very deep networks by avoiding the problem of gradient disappearance. • a ResNeXt that extends the idea of ResNet by introducing a modular structure with multiple “paths” per residual block allowing for better feature representation, thus improving model performance. It can be more efficient than ResNet in some situations. • a VGG (for “Visual Geometry Group” in English) known for its simple and deep architecture, with small convolution layers (3x3) and whose simplicity facilitates its training and which has shown good performance on various data sets; • an EfficientNet that offers a balance between network depth, layer width, and image resolution using a scaling factor, so as to be more resource efficient, offering good performance with a relatively small number of parameters. • an Inception module (like GoogLeNet) that uses inception modules that combine convolutions of different sizes to capture information at different scales, providing efficient feature extraction at different resolutions, which improves the model's ability to learn complex patterns.
[0092] In one embodiment, two identical foundation models may be used in parallel or a Siamese neural network, preferably composed of two identical convolutional neural networks. Advantageously, this will subsequently make it possible to process two images in parallel as a reference retinal image and a retinal image to be authenticated.
[0093] In one embodiment, the feature vectors generated by the foundation model (and therefore by the refined model) have a dimension greater than or equal to 5. Having high-dimensional feature vectors has several advantages such as: - a rich representation of the data, the feature vector can indeed capture a wider range of variations and details in the input data. This allows the model to encapsulate richer and more complex information about the characteristics of the data. - better sensitivity in identification: in fact, in the case of authentication or identification of an individual, having high-dimensional characteristic vectors offers more freedom to organize the data in a discriminative manner. - robustness to minor data transformations and perturbations. This can help improve the generalization of the model to test data or to data of a similar but slightly different nature. - ease of adaptation, high-dimensional feature vectors can be more flexible and adaptable to different tasks or different datasets. The foundation model can then be more easily adjusted / refined (called “fine-tuned” in English) for specific tasks without losing accuracy or representativeness. - complex encoding capability by allowing the model to capture complex patterns and relationships between data, which is essential for tasks such as authentication or identification of an individual.
[0094] It should be noted that the dimension should not be too large either in order to avoid overfitting, so a compromise must be found.
[0095] For example and in a non-limiting manner, the characteristic vector can reside in a metric space of a dimension equal to 127 or 1024 or even 2048 as is possible with a ResNet50 foundation model.
[0096] Advantageously, the generated characteristic vectors can be normalized.
[0097] The metric space can be, for example: a Euclidean space equipped with the Euclidean distance, a hypersphere equipped with either the Euclidean distance or the length of the smallest arc or a Riemannian manifold equipped with a metric tensor which can be modeled by a neural network.
[0098] In one example, the chosen foundation model is a ResNet50 v2 (i.e. second generation) projecting images onto a hypersphere of dimension 1023 or 511 provided with a classical Euclidean distance.
[0099] Step S20 relates to the reception of the first set of raw retinal images El. As illustrated in [Fig.3], the first set of raw retinal images El comprises a single raw retinal image Im for each subject and therefore a single image per identity ID (each subject having an identity per eye / retina).
[0100] These raw retinal images are preferably color images, for example, in RGB format (HxWx3, where H represents the height of the image in pixels, W the width of the image in pixels and the 3 corresponds to the 3 channels Red, Green, Blue). These received raw retinal images can also be in gray level (HxWxl). In this case, they can be artificially transformed into an RGB image by concatenating the same image three times during a pre-processing (see step S25).
[0101] Advantageously, public databases can be used such as “Eyepacs”, “Messidor”, “DDR”, “Aptos”, “ODIR5K” and / or “REFUGE”, so as to have a sufficiently large volume of images to pre-train the model. foundation, ranging from a few thousand raw retinal images to over 120,000.
[0102] Step S30 relates to the reception of the second set of raw retinal images E2. As illustrated in [Fig.4], the second set of raw retinal images E2 comprises at least two raw retinal images Im of the same eye for each subject, therefore by identity ID.
[0103] In the same way as in step S20, the raw retinal images can be color or black and white images, and undergo the same pre-processing (see step S35).
[0104] Advantageously, public databases can be used such as “jichiDR” and / or “yangxi”, so as to have a sufficiently large volume of images to train / refine the pre-trained foundation model (also called “fine-tuning” in English). The number of images in this second set E2 can, for example, vary from a few thousand to around ten thousand raw retinal images.
[0105] Step S40 relates to the generation of sister retinal images of the first set of images EL. The foundation model requiring a very large number of images for its pre-training while the number of raw retinal images available is often limited, an advantageous solution consists of generating sister retinal images from the images of the first set El as illustrated in [Fig.3].
[0106] Thus, for each raw retinal image of the first set of images El, it is possible to create at least one sister retinal image by applying at least one image processing taken from: a change in color, an elastic deformation, a change in brightness, a degradation of the image.
[0107] Each sister retinal image is annotated with the identity of the subject from whom the raw retinal image originates and can also be annotated with the list of image processing applied. This will allow during pre-training to define similarity relationships between sister retinal images.
[0108] Advantageously, at least two, preferably at least ten, sister retinal images are generated from each raw retinal image, in order to have a substantial number of images for the first training database B1.
[0109] Description of the different applicable image processing operations:
[0110] Color change (step T10)
[0111] Advantageously, the color change comprises at least one of: a histogram correspondence between the retinal image and a reference retinal image chosen randomly from a set of reference retinal images, a random modification of the channels of the retinal image in a color space (for example: the “Red, Green, Blue” or RGB space, the “Cyan, Magenta, Yellow" or CMY which can be supplemented by an additional black or K component, the "Hue, Saturation, Value" or HSV space, or the "Luminance, a, b" or LAB space), a conversion of the retinal image into black and white (i.e. into grayscale).
[0112] In the case of histogram matching between channels of an image and channels of a reference image (as illustrated in [Fig.7]), a selection of about a hundred images from the dataset such as Eyepacs, showing an exotic variety of fundus retinal colors as illustrated in [Fig.6], can, for example, be used to randomly choose reference images.
[0113] The overall motivation for the color augmentations is to avoid, for ethical reasons, that the model uses retinal skin pigmentation as a discriminating feature and instead focuses on veins, arteries and vascular network structure. Another beneficial effect of this augmentation is the local modification of brightness and saturation within the images, potential biases for the model during the unsupervised prior self-training (step S60).
[0114] A grayscale conversion can also be applied after histogram matching, for example to 10% of the images, for better generalization of the model to grayscale images.
[0115] Elastic deformation (step T20)
[0116] Advantageously, the elastic deformation is obtained by the application of a predefined displacement field chosen randomly from a set of predefined displacement fields.
[0117] Preferably, the elastic deformation is of the diffeomorphic type in order to simulate small changes in the structure of the vascular network which may occur due to material parameters, acquisition fields, diseases, treatments or simply aging. This increase is particularly interesting because it improves the robustness of the pre-trained foundation model against distortions of the biometric signature.
[0118] For this, the deformation problem can be formulated as an ordinary differential equation (ODE) parameterized by a stationary velocity field (SVF):
[0119]
[0120] [Equation 1] 3 <j) Vte [0; l],^=v(([)t);([>0 = Id
[0121] The deformation field ¢, deforming the image, is obtained by integrating the ordinary differential equation:
[0122] [Equation 2] [01231 ¢ = ^ = 111+(,,-(^) dt
[0124] Provided that the stationary velocity field v is continuously differentiable once in the image domain, the deformation field is diffeomorphic, a property preserving the coherence of structures within the retinal image because it will not generate abrupt or physically impossible transformations such as tears, overlaps or blends. Moreover, the nature of the deformation field is controlled by the shape of the stationary velocity field which, for example, can be sampled from 8 predefined stationary velocity field patterns with random scales, as illustrated in [Fig.8], as well as a corresponding example result in [Fig.9].
[0125] Change of brightness (step T30)
[0126] The change in brightness is intended to simulate the effects on retinal images of variations in illumination induced by environments or equipment with poor lighting, individuals who move during acquisition of the retinal image or pupils that are not sufficiently dilated, as illustrated in [Fig. 10].
[0127] For example, this change in brightness can be based on a logarithmic image processing model (called "Logarithmic Image Processing" in English), adapted to images acquired in transmitted light, which consists of multiplying / adding in a logarithmic manner the image with an illumination map sampled from 3 predefined models with random parameters.
[0128] [Fig.11a] and [Fig. 11b] represent two examples of illuminance maps modeling illuminance variations. The first map ([Fig.11a]), on which a random rotation is applied, gives an illuminance drift when it is multiplied logarithmically with each channel of an image while the other map ([Fig. 11b]), with a randomly positioned Gaussian point, gives a darkening point after logarithmically adding to each channel.
[0129] [Fig. 12] gives an example of sister retinal images obtained via different brightness changes applied to the same original retinal image located at the top left (NB: the original retinal image may be a raw retinal image or may have already undergone image processing and / or image pre-processing).
[0130] Image degradation (step T40)
[0131] Image degradation includes changes that degrade the quality of acquisitions, such as noise, blur or desaturation.
[0132] For example, it is possible to add random noise at the pixel level, sampled uniformly between [-1, 1] and scaled to 10% of the maximum value (here 255). For adding blur, it is possible to apply a simple average filter to the image of the maximum value (here 255). As for desaturation, it can be achieved by multiplying the saturation channel of the image in the Hue, Saturation, Lightness (HSL) color space by a predefined desaturation halo with a random scale.
[0133] An example of sister images obtained with the different degradations mentioned applied to an original retinal image (top image) is shown [Fig. 13] (NB: the original retinal image may be a raw retinal image or one that has already undergone image processing and / or image pre-processing).
[0134] Translation (step T50)
[0135] Advantageously, a translation can also be applied to the sister images. This makes it possible to eliminate part of the bias caused by the relative position of the optic disc (also called optic nerve) or the acquisition field of the initial image.
[0136] As illustrated in [Fig.14], this translation consists of translating the optical disc horizontally (Ax) and / or vertically (Ay) in the image as, for example, (Ax, Ay) e [-aODx, a(S - ODx)] x [-aODy, a(S - ODy)] where 0 < a < 1 would be a factor controlling the extent of the translation along the image axis.
[0137] It should be noted that it is thanks to the application of these different image processing operations that it was possible to aggregate different public databases of retinal images to create training databases that are sufficiently substantial to pre-train and train the different models used, thereby successfully increasing the raw retinal images in a plausible manner, particularly in terms of color, illumination and retinal geometry.
[0138] Order of application of the different image treatments
[0139] If a sister image is generated by applying more than one of the image processes previously described, the order of application can be any. In one embodiment, the order defined in [Fig.5] is followed. It has been observed that implementing this order in the implementation of the processes allows for subsequent performance improvements during the various trainings of the foundation model. In particular, it is particularly advantageous to apply the colorimetry change before the brightness change.
[0140] As stated previously, the method for obtaining a refined model comprises, among other things, step S50 relating to the generation of the first training database B1. This step consists of constituting the first training database B1 so as to comprise the images of the first set of images El plus the sister images previously generated from this first set of El images.
[0141] In step S60, the pre-training of the foundation model on the first database B1 is carried out in a self-supervised manner by defining similarity relationships between the sister retinal images. During this self-supervised pre-training, the foundation model creates supervised learning tasks from the data available in the first training database B1, then learns to solve them.
[0142] In one embodiment, this pre-training is achieved by also defining dissimilarity relationships between the non-sister retinal images.
[0143] The foundation model, which in one example is a neural network, can be trained with a stochastic gradient descent type optimization algorithm or the Adam Solver type algorithm.
[0144] This pre-training is carried out in such a way as to minimize a cost function (called “loss function” in English) which can be: - a non-contrastive cost function (i.e., defining similarity relations between sister retinal images), such that: • the “Barlow twins loss” which encourages the similarity of representations of two images of the same identity (i.e., two sister images) while discouraging redundancy by using a correlation matrix to achieve this objective; • the “VICReg loss” which optimizes intra-class variability while maintaining correlation between images, which uses variance-based regularization; • “BYOL loss” (for “Build Your Own Latent” in English) with an asymmetric architecture which maximizes the correlation between two encoders while using an asymmetric architecture; or - a contrastive cost function using positive and negative image samples (i.e., also defining similarity relations between sister retinal images and dissimilarity between non-sister retinal images), such that: • the triplet loss function which minimizes the distance between an anchor image and a positive image while maximizing the distance with a negative image; • the “InfoNCE loss” (for “Noise Contrastive Estimation” in English) which measures the similarity between a pair of samples in maximizing the ratio of the probability of a positive pair to the sum of the probabilities of the positive and negative pairs; • the “NT-Xent loss” (for “Normalized Temperature-Scaled Cross-Entropy Loss”) which extends the principle of contrastive noise estimation (called “Noise-Contrastive Estimation” or NCE in English) used to measure the probability that a pair of samples comes from the same distribution, by normalizing the scores by a temperature. • the “m-repulse-only loss” which is used in the context of triplet loss, and introduces a repulsion component aimed at increasing the separation between negative samples.
[0145] The concept of triplet loss is based on the idea of example triplets, which consist of three samples: an anchor, a positive example, and a negative example. The goal is to learn representations such that the distance between the anchor and the positive example is minimized, while the distance between the anchor and the negative example is maximized, with a defined margin. To illustrate this triplet idea, the anchor could be a raw retinal image associated with one identity, the positive example would be a sister image from that same raw retinal image, and the negative example would be a raw or sister image from another identity.
[0146] In one example, the formula used for this “triplet loss” is as follows:
[0147] [Equation 3]
[0148] L(xæ xp, xn) = max( (0, ||fw(xa) -fw(xp) ll^' llfw(xa) -tw(xn) ||m))
[0149] where / w is a neural network trained with stochastic gradient descent, xa is the anchor, xp is the positive example and is sampled from the same identity as the anchor, and xn is the negative example and is sampled from a different identity than the anchor. As illustrated in [Fig. 15], the live triplet extraction strategy consists of selecting from mini-batches the hardest triplet for each anchor by mitigating them with randomly selected semi-hard triplets, with the hardest triplets satisfying the following equation:
[0150] [Equation 4]
[0151] ||fw(xj-fw(xp)||2<||fw(xa)-fw(xn)|^< HW**)-%(xp) \\\ + m
[0152] for faster convergence, better penalization and at the same time to avoid bad local minima in the early stages.
[0153] In one embodiment, the pre-training can be done only from pairs of sister images having received at least one image processing, so that:
[0154] [Equation 5]
[0155] Xa = ta (x.) > Xp = tp (x.), Xn = tn (Xj), xr xr x? (ta, tp, tn) ~ T3
[0156] in which and %j are sets of retinal images associated with subjects z and j, and r represents the set of all possible augmentations.
[0157] Concerning step S70 relating to the generation of the sister retinal images of the second set of images, in the same manner as in step S40 and for each raw retinal image of the second set of images E2, it is possible to create at least one sister retinal image by applying at least one image processing taken from: a change of color, an elastic deformation, a change of brightness, a degradation of the image, as described previously. Only the application of the image translation is not recommended for generating the sister images which will be used for the generation of the second database because this degrades the performance of the refined model instead of improving it as was the case during the pre-training of the foundation model.
[0158] In one embodiment, each sister retinal image is annotated with the identity of the subject from which the raw retinal image originated and may also be annotated with the list of image processing applied. This allows for the selection of sets of sister and cousin images during training / refinement.
[0159] Advantageously, at least ten sister retinal images are generated from each raw retinal image, in order to have a substantial number of images for the second training database B2.
[0160] Step S80 consists of constituting the second training database B2 comprising the raw retinal images of said second set and the sister images previously generated from the second set of images E2. In this way the second training database B2 comprises, for each identity: - the at least two raw retinal images; - at least one sister image for each of the at least two raw retinal images, so as to have at least one pair of non-sister retinal images sharing an identity (i.e., at least one pair of cousin images).
[0161] [Fig.3] and [Fig.4] illustrate the different relationships that can exist between the retinal images. As defined previously, the images formed from the same raw retinal image are called sister retinal images. The original raw retinal image can also be considered as being a sister retinal image of the other sister retinal images generated from it, the only difference being that it has not undergone any image processing, except possibly the image pre-processing (see steps S25 and S35).
[0162] Two retinal images are defined as cousins when these images have been generated from two raw retinal images from the same identity. In the same way, two raw retinal images sharing the same identity (i.e. from the same eye of the same subject or individual) are said to be cousins.
[0163] During step S90, obtaining the refined model is done by training the pre-trained foundation model on the second database B2, in a supervised manner and by defining similarity relationships between the cousin retinal images. This operation is called fine-tuning.
[0164] Advantageously, only a part of the pre-trained foundation model can be trained / refined, this can for example be a particular layer of neurons or a separate block.
[0165] The pre-trained foundation model, which in one example is a neural network, can be trained / refined with a Stochastic Gradient Descent type optimization algorithm or the Adam Solver type algorithm.
[0166] In one embodiment, the training is performed by also defining dissimilarity relationships between the non-cousin and non-sister retinal images.
[0167] In one embodiment, during pre-training and refinement, the model (i.e., foundation model and pre-trained foundation model) is penalized with the same cost function. In particular, the same cost functions cited previously for pre-training can be used.
[0168] In an embodiment in which the cost function is a “triplet loss”, the training can be done only from pairs of cousin images having received at least one image processing, so that:
[0169] [Equation 6]
[0170] Xa = ta(xt), xp = tp(x2), xn = tn(x3), (xb x2)
[0171] This then makes it possible to keep the pairs of raw (i.e., non-augmented) retinal images from the second database (instead of needing a third set of raw retinal images from a group of individuals, each individual being identified by at least two raw retinal images acquired on the same eye and therefore having the same identity) to validate the refined model (see steps S100 and S110).
[0172] Step S120 consists of providing the refined model (or the training parameters of the refined model) which can subsequently be used to authenticate or identify an individual from at least one retinal image acquired on one of his eyes. In particular, the refined model (or all of its training parameters) can be stored in a database.
[0173] The refined model thus obtained which, thanks to artificial intelligence for each retinal image received, generates at least one associated characteristic vector makes it possible to avoid having to go through extraction steps, such as segmentation, or recognition of particular characteristics of the image such as the vascular network.
[0174] Advantageously, the method for obtaining a refined model may comprise optional additional steps, such as pre-processing of the received raw retinal images S25 and S35, validation of the refined model SI 10 by receiving a validation database.
[0175] Steps S25 and S35 of pre-processing the received raw retinal images aim to format the raw retinal images in a particular format so as to facilitate their subsequent processing and analysis. For this, at least one image pre-processing can be applied among: a change of color space, a resizing, a cropping, a filtering.
[0176] [Fig. 16a], [Fig. 16b], [Fig. 16c] and [Fig.16d] illustrate an example of applicable image preprocessing where a retina-conscripted square is calculated in order to remove as much unnecessary black background as possible from an acquisition and normalize the image resolution by restricting it only to the retina.
[0177] [Fig. 16a] represents the original raw retinal image of dimension 1944x2592, and [Fig. 16b] represents the image restricted to the square conscripted to the retina of dimension 640x640.
[0178] Two options are then possible: - either the image is directly resized to a working resolution (for example HxW=224x224 or HxW=640x640); - either it is first resized to an intermediate resolution (for example 640x640), the position of the optical disc is determined by a pre-processing algorithm (which may involve a second neural network), then the image is resized to a working resolution and a crop in the shape of a "moon" (or disc) centered on the position of the optical disc is applied to it to keep only the parts of the image rich in vascular networks.
[0179] [Fig. 16c] represents the retinal image restricted to the square conscripted to the retina which has been resized to a dimension of 224x22.4, [Fig.16d] represents the image with a “moon-shaped crop” of dimension 224x140.
[0180] Steps S100 and S110 of validating the parameters of the refined model to evaluate the generalization capacity of the model during the training process, consist of receiving a validation base comprising a third set of raw retinal images from a group of individuals, each individual being identified by at least two raw retinal images acquired on the same eye and therefore having a same identity and which corresponds to step S100. Then, validation step SI 10 is carried out on this validation basis.
[0181] Advantageously, the validation base can be included in the second training database, such that equation 6 becomes:
[0182] [Equation 7]
[0183] xa = X],xp = x2, xn = x3,x{, x2~x?,
[0184] The method described above can advantageously be implemented in a device to obtain a refined model comprising: - at least one input configured to carry out the reception steps S10, S20 and S30, as well as the optional step S100; - at least one processor configured to carry out steps S40 to S90, as well as optional steps S25, S35 and S110; - at least one output configured to perform step S120.
[0185] [Fig.2] illustrates an embodiment of the device for obtaining a refined model 200.
[0186] In this embodiment, the device 200 comprises an input interface 203 for receiving input data and an output interface 204 for providing output data. Examples of input data may be the foundation model and the different sets of raw retinal images. Examples of output data may be the refined model (with or without its training parameters), the sister images generated from the sets of raw retinal images, or the first B1 and the second B2 training databases generated.
[0187] The device 200 may also comprise a memory 201 for storing program instructions loadable into the processor 202 such as the execution of the different steps of the method previously described and illustrated in [Fig.l]. The memory 201 may also store data and information useful for carrying out the steps of the method as described above.
[0188] The processor 202 can be for example: - a processor or processing unit capable of interpreting instructions in a computer language, the processor or processing unit being able to include, be associated with or be attached to a memory comprising the instructions, or - the association of a processor / processing unit and a memory, the processor or processing unit capable of interpreting instructions in a computer language, the memory comprising said instructions, or - an electronic card in which the steps of the invention are described within the silicon, or - a programmable electronic chip such as an FPGA chip (for “Field-Programmable Gate Array”).
[0189] To facilitate interaction with the device 200, a screen 205 and a keyboard 206 may be provided and connected to the processor / computer circuit 202.
[0190] DEVICE AND METHOD FOR AUTHENTICATING AN INDIVIDUAL
[0191] According to another aspect, the present disclosure relates to an authentication device and method for enabling the authentication of an individual from at least one retinal image acquired on at least one eye of said individual and from a refined model obtained using the device or method previously described.
[0192] As illustrated in [Fig.17], the authentication process comprises the following main steps: - step S200: receive at least one retinal image to be authenticated acquired on the individual; - step S210: receiving the refined model configured to, from a received retinal image, generate at least one characteristic vector associated with the received retinal image; - step S220: generating, using the refined model, at least one vector of characteristics from the at least one retinal image to be authenticated; - step S230: calculating at least one estimated similarity score between at least one reference vector associated with the individual and the at least one characteristic vector relating to the retinal image to be authenticated; - step S240: authenticate the at least one retinal image to be authenticated as belonging to the individual from an authentication condition based on said at least one similarity score; - step S250: provide at least one of: the at least one calculated similarity score, a result relating to the authentication of the at least one retinal image to be authenticated, a command relating to an action to be carried out.
[0193] In step S200 of receiving the retinal image to be authenticated, at least one image pre-processing (like those previously defined) can be applied to the at least one retinal image to be authenticated before generating the at least one associated characteristic vector.
[0194] An error message may further be generated and provided if the quality of the received retinal image does not meet certain predefined criteria such as, for example: sharpness, contrast, sufficient surface area of the retina, presence of a certain area of the retina, etc.
[0195] In the context of a simple authentication, a single retinal image of one of the two eyes of the individual can be used. In order to strengthen security, a double authentication can be carried out from at least one retinal image of each eye of the individual to be authenticated.
[0196] Step S210 consists of receiving the refined model previously generated using the method or device according to any of the embodiments previously described, for example from the database in which it was previously stored.
[0197] In step S220, a feature vector is generated by the refined model and, as previously described, the feature vector may have a dimension greater than or equal to 5.
[0198] In step S230, the calculation of the at least one estimated similarity score between at least one reference vector associated with the individual (which can be stored in a memory or a fixed or mobile medium, such as a USB key, included or not in the authentication device, for example) and the at least one vector of characteristics relating to the retinal image to be authenticated, can be done in two different ways: - from the contrastive cost function associated with the refined model; - from at least one distance calculated between the at least one vector of reference associated with the individual and the at least one authentication vector relating to the retinal image to be authenticated, preferably this at least one distance is a Euclidean distance as previously described.
[0199] In one embodiment, during step S240 relating to the authentication of the individual, the authentication condition comprises the comparison of the at least one similarity score previously calculated with at least one predefined threshold. This at least one predefined threshold can be calculated following the application of a cross-validation method carried out on the refined model using a validation base comprising a third set of raw retinal images of a group of individuals, each individual being identified by at least two raw retinal images of the same identity.
[0200] Thus, when the predefined threshold is reached, this means that the retinal image to be authenticated belongs to the individual with whom the at least one reference vector is associated.
[0201] Double authentication is also possible; for this, it is simply necessary to have a retinal image to be authenticated and previously acquired on the second eye of the individual as well as at least a second reference vector associated with this second eye.
[0202] As previously described, the validation base can advantageously be included in the second training database.
[0203] During step S250 relating to the provision of an output and according to the needs of the individual or the system on which the individual wishes to be authenticated (for example: a smartphone or any other object requiring a sort of entry key, a secure application / account / action such as access to an email box, a bank account or the validation of an online purchase), it may be interesting to provide the at least one calculated similarity score, the result relating to the authentication of the at least one retinal image to be authenticated, or even a particular command to be able to carry out an action such as those cited as examples.
[0204] Advantageously, the authentication method may include optional steps such as: - the reception of another refined model identical to the first received so as to function as a Siamese network as previously mentioned: - the reception of at least one previously acquired reference retinal image of the individual; - applying at least one image pre-processing to the at least one received reference retinal image; - the generation, using the refined model, of at least one reference vector associated with the individual from at least one pre-processed or non-pre-processed reference retinal image.
[0205] The authentication method described above can advantageously be implemented in an authentication device comprising: - at least one input configured to carry out the reception steps S200 and S210, as well as the optional step of receiving a reference image; - at least one processor configured to carry out steps S220 to S240 as well as the optional steps of pre-processing the received retinal images (those to be authenticated and / or those of reference) and of generating at least one reference vector; - at least one output configured to perform step S250.
[0206] In the same manner as previously described, [Fig.2] illustrates an embodiment of the device for obtaining a refined model 200.
[0207] Advantageously, the memory 201 can also be configured to store, among other things, the at least one reference vector associated with the individual to be authenticated.
[0208] Advantageously, the authentication device may also comprise at least one acquisition device configured to acquire raw retinal images of the individuals to be authenticated. The at least one acquisition device is chosen from: a non-mydriatic retinal camera, a mydriatic retinal camera, a digital ophthalmoscope, an optical coherence tomograph (called "Optical Coherence Tomography (OCT), a retinal angiograph, a smartphone adapter or even a telemedicine device.
[0209] DEVICE AND METHOD FOR AUTHENTICATING AN INDIVIDUAL
[0210] According to another aspect, the present disclosure relates to an identification device and method for enabling the identification of an individual from at least one retinal image acquired on at least one eye of said individual and from a refined model obtained using the device or method previously described.
[0211] As illustrated in [Fig. 18], the identification method comprises the following main steps: - step S300: receive at least one retinal image acquired on the individual to be identified; - step S310: receiving the refined model configured to, from a received retinal image, generate at least one characteristic vector associated with the received retinal image; - step S320: generating, using the refined model, at least one characteristic vector from the at least one received retinal image; - step S330: calculating at least one estimated similarity score between a set of reference vectors associated with a set of individuals whose identity is known and said at least one characteristic vector relating to the received retinal image; - step S340: determining the identity of said individual from among the set of individuals whose identity is known, from a condition based on said at least one similarity score; - step S350: provide at least one of: said at least one calculated similarity score, a result relating to the identity of said individual associated with said at least one received retinal image, a command relating to an action to be carried out.
[0212] Steps S300 to S320, S340 and S350 are respectively similar to steps S200 to S220, S240 and S250 previously described.
[0213] During step S330 relating to the calculation of at least one similarity score and in the same way as for step S230, this calculation of the at least one similarity score estimated between the set of reference vectors and the at least one vector of characteristics relating to the retinal image of the individual to be identified, can be done in two different ways: - from the contrastive cost function associated with the refined model; - from at least one distance calculated between each vector of the set of reference vectors and the at least one authentication vector relating to the retinal image to be authenticated, preferably this at least one distance is a Euclidean distance as previously described.
[0214] Advantageously, the authentication method may include optional steps such as: - the reception of a set of reference retinal images previously acquired and associated with the set of individuals whose identity is known; - applying at least one image pre-processing to at least one of the received reference retinal images; - the generation, using the refined model, of the set of reference vectors from the set of reference images, pre-processed or not.
[0215] The identification method described above can advantageously be implemented in an identification device comprising: - at least one input configured to carry out the reception steps S300 and S310, as well as the optional step of receiving all of the reference images; - at least one processor configured to carry out steps S320 to S340 as well as the optional steps of pre-processing the received retinal images (those of the individuals to be identified and / or those of reference) and of generating all of the reference vectors; - at least one output configured to perform step S350.
[0216] In the same manner as previously described, [Fig.2] illustrates an embodiment of the device for identifying an individual 200.
[0217] Advantageously, the memory 201 can also be configured to store, among other things, the set of reference vectors.
[0218] Advantageously, the identification device may also comprise at least one acquisition device configured to acquire raw retinal images of the individuals to be identified. COMPUTER PROGRAM PRODUCT
[0219] The present disclosure also relates to a computer program product intended to provide the optimal sequence of actions causing the evolution of a complex system or device from an initial state to a final state, the computer program product comprising instructions which, when the program is executed by a computer, cause the computer to automatically execute the steps of the method for obtaining a refined model, the steps of the method for authenticating an individual and / or the steps of the method for identifying an individual, according to one of the embodiments described above.
[0220] The computer program product for performing the various methods described above may be written in the form of computer programs, code segments, instructions, or any combination thereof, to individually or collectively instruct or configure the processor or computer to operate. as a special-purpose machine or computer for performing the operations performed by the hardware components. In one example, the computer program product includes machine code directly executed by a processor or computer, such as machine code produced by a compiler. In another example, the computer program product includes higher-level code that is executed by a processor or computer using an interpreter. Programmers of ordinary skill in the art can readily write the instructions or software based on the block diagrams and flowcharts illustrated in the drawings and the corresponding descriptions in the specification, which disclose algorithms for performing the operations of the method as described above. COMPUTER-READABLE STORAGE MEDIUM
[0221] The present disclosure also relates to a computer-readable storage medium comprising instructions which, when the program is executed by a computer, cause the computer to perform the steps of the method for obtaining a refined model, the steps of the method for authenticating an individual and / or the steps of the method for identifying an individual, according to one of the embodiments.
[0222] According to one embodiment, the computer-readable storage medium is a non-transitory computer-readable storage medium.
[0223] The computer programs implementing the various methods described in the present embodiments may generally be distributed to users on a computer-readable storage medium, such as, but not limited to, an SD card, an external storage device, a microchip, a flash memory device, a portable hard drive, and software websites. From the distribution medium, the computer programs may be copied to a hard drive or similar intermediate storage medium. The computer programs may be executed by loading the computer instructions, either from their distribution medium or from their intermediate storage medium, into the execution memory of the computer, configuring the computer to act in accordance with the method of the present disclosure.All these operations are well known to people skilled in computer systems.
[0224] The instructions or software for controlling a processor or computer to implement the hardware components and perform the methods described above, and all associated data, data files, and data structures, are recorded, stored, or fixed in or on one or more non-transitory computer-readable storage media. Examples of non-transitory computer-readable storage media include read-only memory (ROM), random access memory (RAM), flash memory, CD-ROMs, CD-Rs, CD-I-Rs, CD-RWs, CD-I-RWs, DVD-ROMs, DVD-Rs, DVD+Rs, DVD+Rs, DVD+ R, DVD+ R, DVD+ R, DVD+ R, DVD+ R, DVD+ R, DVD+ R, etc., DVD- Rs, DVD+ Rs, DVD- RWs, DVD+ RWs, DVD- RAMs, BD- ROMs, BD- Rs, BD- R LTHs, BD- Res, magnetic tapes, floppy disks, magneto-optical data storage devices, optical data storage devices, hard disk drives, solid-state drives, and any device known to a person skilled in the art capable of storing the instructions or software and all associated data, data files and data structures in a non-transitory manner and of providing the instructions or software and all associated data, data files and data structures to a processor or computer so that the processor or computer can execute the instructions.In one example, the instructions or software and all associated data, data files, and data structures are distributed across network-coupled computer systems, such that the instructions and software and all associated data, data files, and data structures are stored, accessed, and executed in a distributed manner by the processor or computer. Examples
[0225] The present invention will be better understood by reading the following example which illustrates the invention in a non-limiting manner.
[0226] Databases used
[0227] For this example, the following public databases were used
[0228] [Tables 1] Public dataset Number of identities Number of images Number of images per identity (average, [min, max]) Eyepacs 88702 88702 1.0; [1,1] messidor2 1748 1748 1.0; [1,1] DDR 13 673 13 673 1.0; [1,1] Aptos 5590 5590 1.0; [1,1] 0RDIR5K 8000 8000 1.0; [1,1] REFUGE 2400 2400 1.0; [1,1] jichiDR 2391 6946 2.9 ; [ 2, 13 ] Yangxi 5561 11 122 2.0 ; [2, 2] JSIEC1000 35 72 2.1; [2,4] AiScreeningsPublic 124 391 3.2; [2, 8] CATARACT 71 144 2.0; [2, 4] Diaretdb 7 14 2.0; [2, 2] DR12 260 612 2.4; [2, 10] Eophta 140 283 2.0; [2, 3] FIRE 39 119 3.1; [2.9] G1020 67 154 2.3; [2, 4] IDRiD 7 14 2.0; [2, 2] longDRscreening 140 552 3.9; [3, 5] ORIGA 4 8 2.0; [2,2] RIDB 20 100 5.0; [5, 5] ROC 1 2 2.0; [2, 2]
[0229] In order to constitute the different sets of raw retinal images E1, E2 and E3 necessary for the implementation of the method to obtain a refined model, some of these databases were aggregated as follows:
[0230] [Tables2] Thursday public data^ Number* of identities;- Number* of images^ Number'of images per-identity -(average *E * fmms*niax])& First -set E1 — £; -messidorS -4--DDR4- - 120090^ 120«Qa Second -set -E2 — • {large • subset of 7652s 22ÆtX13-]D Third- set -E3 — ■ {JSIEC1000-+-- CATARACT-F-tea-- DR12.+-MB-4-FIRE4-GÎ020-+-ÏPBO +-hAmg •+ ORIGA-r ■ RIDBsmall subset -of 'and (-100 -identities -and^ ^■200-identities -respectively}n 1215s 3517s
[0231] Image augmentation / generation of sister images
[0232] In order to artificially increase the number of images to generate the first B1 and the second B2 training databases in order to pre-train the foundation model and then train / fine-tune it, sister retinal images were generated for each raw retinal image of the first El and second E2 sets of images, respecting the order of image processing recommended in [Fig.5].
[0233] Ten sister retinal images were generated for each raw retinal image of the first E1 and second E2 set by randomly applying one to several of the previously defined image processing operations, namely: a color change, an elastic deformation, a brightness change, an image degradation and / or a translation (applicable for the images of the first set only).
[0234] Regarding the color change, a selection of 121 images from the Eyepacs dataset featuring an exotic variety of fundus retinal colors (as shown in [Fig.6]) was used to randomly choose a reference retinal image to apply the histogram matching.
[0235] A grayscale conversion was also performed after applying histogram matching on 10% of the images for better generalization of the model to grayscale images.
[0236] The applied elastic deformation was a diffeomorphic type deformation as defined previously and according to equations 1 and 2.
[0237] The applied brightness change was based on a logarithmic image processing model (called "Logarithmic Image Processing" in English), which consisted of logarithmically multiplying / adding the image with an illumination map sampled from 3 predefined models with random scales and intensities modeling illumination drifts, darkening spots and general darkening of the image.
[0238] Image degradation consisted of adding noise, blur, and / or desaturation. The added noise was random pixel-level noise, sampled uniformly between [-1, 1] and scaled to 10% of the maximum value (here 255). Blurring involved applying a simple average filter to the image. Desaturation was achieved by multiplying the image's saturation channel in the Hue, Saturation, Lightness (HSL) color space by a predefined desaturation halo with a random scale.
[0239] The translation (applicable only for generating sister retinal images on the images of the first set of images El) consisted of translating the optic disc horizontally (Ax) and / or vertically (Ay) in the image such that (Ax, Ay) G [-aODx, a(S - ODx)] x [-aODy, a(S - ODy)] where 0 < a < 1 is a factor controlling the extent of the translation along the image axis.
[0240] Self-supervised pre-training
[0241] The foundation model used is a self-supervised pre-trained fw convolutional neural network on the first database B1 comprising the raw retinal images of the first set of images El and the associated sister retinal images. In particular, the foundation model used is a ResNet50V2 type convolutional neural network whose convolutional layers have been initialized on the ImageNet database, having a normalized L2 projection space of a dimension d = 215 units.
[0242] This pre-training allowed the foundation model to learn a pseudometric pullback between the space of retinal images X Œ jetxwidthx3 and a pre-defined metric eSpace (Sdl, ||.|| ), which represents a hypersphere equipped with the Euclidean norm of the ambient space Rd.
[0243] To carry out this pre-training, image triplets were created from the images of the first training database B1 as follows, in order to apply the contrastive cost function “triplet loss” defined in equation 3:
[0244] Given 2 raw retinal images xt and xj from the first set of images El, and 3 image processings (^ tn) ~ E3 chosen randomly from those defined previously, we define the following triplet of images Xa — ht (xi), Xp — tp (Xj), Xn — tn (Xj), where xa is considered as the anchor, xp is the positive example and xn is the negative example.
[0245] Supervised Training / Refining
[0246] For training the pre-trained foundation model, in order to obtain the refined model, the same contrastive cost function “triplet loss” as defined in equation 3 is used, but this time with the triplet of images defined in equation 5, where xi and x 2 are cousin retinal images (including sharing the same identity) and x 3 is a retinal image from another identity, taken from the second training database B2.
[0247] Results
[0248] Following validation of the refined model on the third set of images E3, the following results are obtained: a false acceptance rate (called “False Acceptance Rate” or FAR in English) of approximately 0.001 and a validation accuracy level (called “Validation Accuracy Level” or VAL in English) of approximately 97%.
Claims
1. Claims Device for obtaining a refined model configured to allow the authentication of an individual from at least one retinal image acquired on an eye of said individual, said device comprising: at least one input configured to receive (S 10, S20, S30): - a foundation model configured to, from a received retinal image, generate at least one feature vector associated with the received retinal image; - a first set of raw retinal images acquired from a plurality of subjects and comprising one raw retinal image per subject; - a second set of raw retinal images acquired from a plurality of subjects, comprising at least two raw retinal images of the same eye for each subject; at least one processor configured to: - for each of the raw retinal images of said first set of images, generating at least one sister retinal image by applying at least one image processing to said raw retinal image and associating with said sister image thus generated the identity of the subject from which said associated raw retinal image originates (S40); - generating a first training database (Bl) comprising the raw retinal images of said first set and their generated sister retinal images (S50); - pre-train the foundation model with the first training database (Bl), in a self-supervised manner, by defining similarity relationships between sister retinal images (S60); - for each of the raw retinal images of said second set of images, generating at least one sister retinal image by applying at least one image processing to said raw retinal image and associating with said sister image thus generated the identity of the subject from which said associated raw retinal image originates (S70); - generating a second training database (B2) comprising the raw retinal images of said second set and their generated sister retinal images, where any pair of images non-sister retinal images sharing an identity of the same subject are said to be cousins (S80); - obtaining said refined model by training on the second training database (B2) said pre-trained foundation model, in a supervised manner, by defining similarity relationships between the cousin retinal images (S90); at least one output configured to provide said refined model (S 120).
2. The device of claim 1, wherein pre-training the foundation model with the first training database (B1), in a self-supervised manner (S60), is performed by also defining dissimilarity relationships between the non-sister retinal images.
3. Device according to claim 1 or 2, wherein obtaining said refined model by training on the second training database (B2) said pre-trained foundation model, in a supervised manner (S90), is carried out by also defining dissimilarity relationships between the non-cousin and non-sister retinal images.
4. The device of claim 3, wherein the training of the pre-trained foundation model (S90) on the second training database (B2) is also done so as to minimize a contrastive cost function from the similarity relationships defined between the cousin retinal images and the dissimilarity relationships defined between the non-cousin retinal images.
5. Device according to one of claims 1 to 4, wherein said at least one image processing applied to generate the sister retinal images comprises a color change (T 10), said color change comprising at least one of: a histogram matching between the retinal image and a reference retinal image randomly chosen from a set of reference retinal images, a random modification of the channels of the retinal image in a color space, a conversion of the retinal image to black and white.
6. Device according to one of claims 1 to 5, wherein said at least one image processing applied to generate the sister retinal images comprises an elastic deformation of the retinal image (T20) by the application of a displacement field predefined randomly chosen from a set of predefined displacement fields.
7. Device according to any one of claims 1 to 6, wherein said at least one image processing applied to generate the sister retinal images comprises a brightness change (T30) comprising at least one of: an overall darkening of the retinal image, a local overexposure of the retinal image, an addition of a darkening spot.
8. Device according to any one of claims 1 to 7, wherein said at least one image processing applied to generate the sister retinal images comprises a degradation of the retinal image (T40) by the addition or removal of at least one of: noise, blurring, desaturation.
9. Device according to any one of claims 1 to 8, wherein at least one of the sister retinal images generated from the raw retinal images of said first set of images is generated by applying at least one translation of the retinal image (T50).
10. Authentication device for enabling the authentication of an individual from at least one retinal image acquired on at least one eye of said individual and from a refined model obtained according to any one of claims 1 to 9, said device comprising: at least one input configured to receive (S200, S210): - at least one retinal image to be authenticated acquired on said individual; - said refined model configured to, from a received retinal image, generate at least one characteristic vector associated with the received retinal image; at least one processor configured to: - generate, using said refined model, at least one characteristic vector from said at least one retinal image to be authenticated (S220); - calculating at least one estimated similarity score between at least one reference vector associated with said individual and said at least one characteristic vector relating to the retinal image to be authenticated (S230);- authenticating said at least one retinal image to be authenticated as belonging to said individual based on a condition; authentication based on said at least one similarity score (S240); at least one output configured to provide as output (S250) at least one of: said at least one calculated similarity score, a result relating to the authentication of said at least one retinal image to be authenticated, a command relating to an action to be performed.
11. Authentication device according to claim 10, wherein the at least one input is also configured to receive at least one reference retinal image of said individual previously acquired, and the at least one processor is further configured to generate, using said refined model, the at least one reference vector associated with said individual from said at least one reference retinal image.
12. Authentication device according to claim 10 or 11, wherein said at least one similarity score is estimated from said contrastive cost function of said refined model.
13. Authentication device according to one of claims 10 to 12, wherein said at least one similarity score is calculated from at least one distance calculated between at least one reference vector associated with said individual and said at least one authentication vector relating to the retinal image to be authenticated, preferably said at least one distance is a Euclidean distance.
14. Authentication device according to any one of claims 10 to 13, wherein the authentication condition comprises the comparison of said at least one similarity score with at least one predefined threshold, said at least one predefined threshold being calculated following the application of a cross-validation method carried out on said refined model using a validation base comprising a third set of raw retinal images of a group of individuals, each individual being identified by at least two raw retinal images of the same identity.
15. Device for identifying an individual from at least one retinal image acquired on at least one eye of said individual, from a refined model obtained according to any one of claims 1 to 9, said device comprising: at least one input configured to receive:
16. - at least one retinal image acquired on said individual to be identified; - said refined model configured to, from a received retinal image, generate at least one vector of characteristics associated with the received retinal image; at least one processor configured to: - generating, using said refined model, at least one characteristic vector from said at least one received retinal image; - calculate at least one estimated similarity score between a set of reference vectors associated with a set of individuals whose identity is known and said at least one vector of characteristics relating to the received retinal image; - determining the identity of said individual from among the set of individuals whose identity is known, from a condition based on said at least one similarity score; at least one output configured to output at least one of: said at least one calculated similarity score, a result relating to the identity of said individual associated with said at least one received retinal image, a command relating to an action to be performed. Computer-implemented method for obtaining a refined model configured to enable the authentication of an individual from at least one retinal image acquired on an eye of said individual, said method comprising: - receiving a foundation model configured to, from a received retinal image, generate at least one characteristic vector associated with the received retinal image (S 10); - receiving a first set of raw retinal images acquired from a plurality of subjects and comprising one raw retinal image per subject (S20); - receiving a second set of raw retinal images acquired from a plurality of subjects, comprising at least two raw retinal images of the same eye for each subject (S30); - for each of the raw retinal images of said first set of images, generating at least one sister retinal image by applying at least one image processing to said raw retinal image and associating with said sister image thus generated the identity of the subject from which said associated raw retinal image originates (S40); - generating a first training database (B1) comprising the raw retinal images of said first set and their generated sister retinal images (S50); - pre-training the foundation model with the first training database (B1) in a self-supervised manner by defining similarity relationships between the sister retinal images (S60); - for each of the raw retinal images of said second set of images, generating at least one sister retinal image by applying at least one image processing to said raw retinal image and associating with said sister image thus generated the identity of the subject from which said associated raw retinal image originates (S70); - generating a second training database (B2) comprising the raw retinal images of said second set and their generated sister retinal images, where any pair of non-sister retinal images sharing an identity of the same subject are said to be cousins (S80);- obtaining said refined model by training on the second training database (B2) said pre-trained foundation model, in a supervised manner, by defining similarity relationships between the cousin retinal images (S90); - providing said refined model (S 120).;
17. Method for authenticating an individual from at least one retinal image acquired on at least one eye of said individual and from a refined model obtained according to the method of claim 16, said method being implemented by computer and comprising: - receiving at least one retinal image to be authenticated acquired on said individual (S200); - receiving said refined model configured to, from a received retinal image, generate at least one characteristic vector associated with the received retinal image (S210); - generating, using said refined model, at least one characteristic vector from said at least one retinal image to be authenticated (S220); - calculating at least one estimated similarity score between at least one reference vector associated with said individual and said at least one characteristic vector relating to the retinal image to be authenticated (S230); - authenticating said at least one retinal image to be authenticated as belonging to said individual from an authentication condition based on said at least one similarity score (S240); - providing at least one of: said at least one calculated similarity score, a result relating to the authentication of said at least one retinal image to be authenticated, a command relating to an action to be performed (S250).
18. A computer-implemented method for identifying an individual from at least one retinal image acquired on at least one eye of said individual and a refined model obtained according to the method of claim 16, said method comprising: - receiving at least one retinal image acquired on said individual to be identified (300); - receiving said refined model configured to, from a received retinal image, generate at least one feature vector associated with the received retinal image (310); - generating, using said refined model, at least one feature vector from said at least one received retinal image (320); calculating at least one estimated similarity score between a set of reference vectors associated with a set of individuals whose identity is known and said at least one feature vector relating to the received retinal image (S330);- determining the identity of said individual from among the set of individuals whose identity is known, from a condition based on said at least one similarity score (S340); - providing at least one of: said at least one calculated similarity score, a result relating to the identity of said individual associated with said at least one received retinal image, a command relating to an action to be performed (S350).;