Improved ophthalmic apparatus
The ophthalmic apparatus uses a detection unit with multiple cameras and an AI module to accurately center ocular inspection devices on the pupil, addressing positioning challenges and enhancing reliability and simplicity in manufacturing.
Patent Information
- Application Number
- PCT/EP2025/066878
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-21
- Filing Date
- 2025-06-17
- Publication Date
- 2025-12-26
AI Technical Summary
Existing ophthalmic apparatus face challenges in accurately and automatically centering ocular inspection devices with respect to the pupil, particularly in the presence of make-up, spurious light sources, shadow areas, partial occlusions, nevi, and tattoos, and when the pupil is outside the camera's field-of-view, requiring complex calibration and lacking sufficient accuracy and operational reliability.
An ophthalmic apparatus equipped with a detection unit comprising multiple cameras and an artificial intelligence module that processes images to identify and predict the pupil's position, even when partially hidden or outside the field-of-view, using a trained convolutional neural network to generate precise position data for the ocular inspection device's actuation unit to achieve continuous centering.
The apparatus ensures high accuracy and operational reliability in maintaining optimal centering of the ocular inspection device, simplifying manufacturing and reducing costs while handling various occlusion and lighting conditions.
Smart Images

Figure EP2025066878_26122025_PF_FP_ABST
Abstract
Description
[0001] IMPROVED OPHTHALMIC APPARATUS
[0002] DESCRIPTION
[0003] The present invention relates to the field of ophthalmic apparatus. In particular, the present invention relates to an ophthalmic apparatus provided with improved means for automatically centering an ocular inspection device with respect to the pupil of the eye of the patient being observed.
[0004] In ophthalmic apparatus, the use of ocular inspection devices, such as the so-called “fundus cameras”, FAP (Fundus Automated Perimeter) devices or OCT (optical coherence tomography) inspection devices is widely known.
[0005] As is known, to perform a correct eye examination, it is important that the optical axis of the ocular inspection device passes through the center of the pupil of the eye being examined and that the ocular inspection device is at the optimal distance from the pupil.
[0006] It is therefore necessary to obtain information about the relative positioning in space between the pupil of the eye being observed and the ocular inspection device. Once processed, this information may be used to move the ocular inspection device in space and bring it to the optimal centering position with respect to the pupil.
[0007] Patent applications EP1442698A1 and US2005117115A1 describe ophthalmic apparatus provided with centering means that allow the position of the pupil to be measured along the optical axis of a fundus camera. Such centering means include special illumination devices to project infrared light towards the cornea, a camera to observe the illuminated anterior segment of the eye and an image processing unit to identify the reflections of the infrared light on the image of the cornea (in the form of light spots) and estimate the axial position of the pupil by measuring the distance between these light spots.
[0008] Patent application WO2022117646A1 describes an ophthalmic apparatus having centering means that include an artificial intelligence unit for detecting the relative position of the pupil, with respect to a fundus camera, from an image of the eye. Manual adjustment means are provided to align a fundus camera with the pupil of the eye based on information provided by the artificial intelligence unit, made available to the user through an appropriate graphical interface.
[0009] Patent application EP04199804A1 describes an ophthalmic apparatus provided with centering means including an artificial intelligence unit to analyze an image of the eye and verify whether or not the pupil is aligned with a fundus camera.
[0010] The ophthalmic apparatus currently available still has some aspects that need improvement. Many ophthalmic devices do not allow for continuous monitoring of the pupil position and require complicated calibration and adjustment procedures to be performed in order to set up the centering means of the on-board ocular inspection device.
[0011] Furthermore, even the most modern ophthalmic apparatus does not yet allow for fully satisfactory levels of accuracy to be obtained with regard to the automatic centering of the ocular inspection device with the pupil, in particular when the imaging of the patient's pupil by means of a video camera is carried out in particular conditions, for example in the presence of make-up on the patient's face, spurious light sources, shadow areas, partial occlusions of the eye, nevi, skin blemishes, tattoos, and so on. Moreover, even the most modern ophthalmic apparatus is not able to provide an automatic centering of the ocular inspection device with the pupil when it is outside the field-of-view of the cameras.
[0012] The main task of the present invention is to provide an ophthalmic apparatus which allows the drawbacks of the prior art, highlighted above, to be overcome or mitigated.
[0013] Within this task, a purpose of the present invention is to provide an ophthalmic apparatus in which it is possible to continuously evaluate the position of the pupil with great accuracy, so as to maintain an optimal centering position of the device during the ocular examination.
[0014] A further purpose of the present invention is to provide an ophthalmic apparatus which is of high operational reliability.
[0015] A further purpose of the present invention is to provide an ophthalmic apparatus that is relatively simple to manufacture on an industrial level, at competitive costs.
[0016] This task and these purposes, as well as other goals which will be apparent from the following description and the accompanying drawings, are achieved, according to the invention, by an ophthalmic apparatus, according to claim 1, proposed below, and the related dependent claims. As general definition, the ophthalmic apparatus, according to the invention, comprises an ocular inspection device and a centering system adapted to position the ocular inspection device in a centering position with respect to the pupil of an eye of the patient being examined.
[0017] The centering system of the ophthalmic apparatus comprises a detection unit including at least one camera adapted to frame the pupil of the eye. The above-mentioned detection unit is adapted to provide a plurality of images of the eye taken from different angles with respect to the pupil at a certain detection instant.
[0018] Preferably, the detection unit comprises a plurality of cameras with different optical axes and oriented so that the pupil is framed by said cameras.
[0019] Preferably, the detection unit comprises a pair of cameras and is configured to provide a pair of images of the eye taken from different angles with respect to the pupil. According to other embodiments of the invention, however, the detection unit may comprise a single camera movable with respect to said ocular inspection device.
[0020] The centering system of the ophthalmic apparatus comprises a data processing unit capable of acquiring the images provided by the detection unit and generating, based on the acquired images, position data indicative of a first relative position occupied by the pupil with respect to the ocular inspection device at a certain detection instant.
[0021] The ophthalmic apparatus centering system comprises an actuation unit adapted to receive and process the position data provided by the data processing unit and to move, based on the acquired position data, the ocular inspection device in space with respect to the pupil so that, after the movement of the ocular inspection device, the pupil is in a second relative position with respect to said ocular inspection device at which the ocular inspection device is centered with respect to the pupil.
[0022] According to the invention, the data processing unit comprises an artificial intelligence module adapted to process the images provided by the detection unit and provide information indicative of a position occupied by the pupil of the eye of the patient being observed.
[0023] In operation, when it has been trained, the artificial intelligence module is capable of identifying the position of the pupil based on visual patterns present in the images provided by the detection unit, both when the pupil is fully shown in said images and when the pupil is partially or fully hidden in said images, for example due to the presence of occlusions when the one or more cameras of the detection unit framed said pupil.
[0024] In operation, when it has been trained, the artificial intelligence module is further capable of predicting the position of the pupil based on visual patterns of the images provided by the detection unit when the pupil is not shown in said images as it is outside the field-of-view of said cameras.
[0025] Preferably, the information provided by said data processing unit include one or more among:
[0026] - one or more semantic segmentation maps of the eye images; and
[0027] - position data indicative of a position occupied by the pupil in a two-dimensional reference system; and
[0028] - position data indicative of a relative three-dimensional position occupied by the pupil with respect to said ocular inspection device.
[0029] In a further aspect, the present invention relates to a method for providing an artificial intelligent module for processing images of an eye provided by a detection unit of an ophthalmic apparatus. In operation, said artificial intelligence module provides information indicative of a position occupied by the pupil.
[0030] In operation, said artificial intelligence module is capable of identifying the position of the pupil based on visual patterns present in said images both when the pupil is fully shown in said images and when the pupil is at least partially hidden in said images.
[0031] In operation, said artificial intelligence module is capable of predicting the position of the pupil based on visual patterns present in said images when the pupil is not fully shown in said images as the pupil is at least partially outside the field-of-view of a camera capturing the eye of a patient.
[0032] The method, according to the invention, includes the following steps:
[0033] - defining the architecture and the hyperparameters of said artificial intelligence module;
[0034] - preparing a dataset for learning the trainable parameters of said artificial intelligence module. Said dataset includes a training dataset, a validation dataset, and a test dataset;
[0035] The method of the invention further comprises the step of carrying out one or more training cycles of said artificial intelligence module by using said training dataset.
[0036] The method further comprises the step of modifying, at one or more of said training cycles, one or more training images of said training dataset.
[0037] Training images can be modified by means of standard data augmentation techniques, which may include, for example, random horizontal flip, rotation, brightness and contrast adjustment. A particularly important aspect of the method of the invention, however, resides in that the training images are modified to simulate one or more circumstances in which:
[0038] - the pupil is at least partially hidden by occlusions; or
[0039] - the pupil is not included or only partially included in the field of view of a camera capturing the eye of a patient.
[0040] The method of the invention further includes the step of validating said artificial intelligence module by using said validation dataset during the training of such artificial intelligence module.
[0041] The method further comprises the step of testing said artificial intelligence module by using said testing dataset, once said artificial intelligence module has been trained.
[0042] Preferably, the above-mentioned dataset preparation step includes acquiring training images of an eye of a patient by means of an ocular inspection device, wherein said training images are acquired in different observation setups
[0043] Said training images are acquired while moving the ocular inspection device with respect to the patient’s eye, so that the training images show the eye in different positions with respect to the ocular inspection device. Moreover, during the acquisition of the training images, while the ocular inspection device is moved, the patient simultaneously moves the eye in different directions, so that the training images show the pupil in different positions with respect to the ocular inspection device.
[0044] Preferably, the above-mentioned dataset preparation step includes labelling the acquired training images and forming a plurality of data groups, each data group including at least a training image and control information associated with said at least a training image.
[0045] Preferably, the above-mentioned dataset preparation step includes checking the control information included in said data groups and discarding control information not complying with predefined quality requirements.
[0046] Preferably, the above-mentioned dataset preparation step includes partitioning the data groups to obtain said training dataset, said validation dataset and said test dataset.
[0047] Preferably, the above-mentioned dataset preparation step includes selecting the acquired training images by discarding duplicated or overly similar training images acquired in said acquisition step.
[0048] Preferably, the above-mentioned control information includes one or more among:
[0049] - one or more semantic segmentation maps of eye images; and
[0050] - position data indicative of a position occupied by the pupil in a two-dimensional reference system; and
[0051] - position data indicative of a relative position occupied by the pupil with respect to said ocular inspection device.
[0052] Further features and advantages of the invention will be better understood with reference to the description given below and to the accompanying figures, provided for merely illustrative and non-limiting purposes, in which:
[0053] - Figures 1, 1A schematically illustrate a perspective view of the ophthalmic apparatus, according to the invention; and
[0054] Figures 2A, 2B illustrate block diagrams of the ophthalmic apparatus, according to the invention, according to some possible embodiments; and
[0055] - Figures 3-5 schematically illustrate some steps of the operation of the ophthalmic apparatus, according to the invention;
[0056] - Figures 6-10, 10A, 11-13 schematically illustrate a method for providing an artificial intelligence module to be included in the ophthalmic apparatus, according to the invention.
[0057] With reference to the aforementioned figures, the present invention relates to an ophthalmic apparatus 1.
[0058] The ophthalmic apparatus 1 preferably comprises a support base 11 and a support structure 12 integrally associated with the base 11 and shaped in such a way as to facilitate the support of the patient's face during the ocular examination.
[0059] The ophthalmic apparatus 1 comprises an ocular inspection device 10 movable in a three- dimensional space with respect to the base 11 and the support structure 12.
[0060] The ocular inspection device 10 preferably consists of a fundus device, such as a fundus camera, a FAP device, or an OCT inspection device.
[0061] Preferably, the ocular inspection device 10 comprises a lens 2 having an optical axis a. During operation of the ophthalmic apparatus (i.e., during the ocular examination), the patient has his / her face supported by the support structure 12.
[0062] The pupil EP of the eye E (or rather, the center of the pupil itself) is thus found in a certain nominal position PN in a fixed three-dimensional reference system (X, Y, Z), for example integral with the support base 11 of the ophthalmic apparatus (Figures 2A-2B).
[0063] Ideally, the nominal position PN of the pupil EP is predefined. However, it may change during the eye examination, due to movements of the patient's eye.
[0064] The ophthalmic apparatus 1 comprises a centering system 3 for positioning the ocular inspection device 10 in a centering position with respect to the pupil EP of the eye E of the patient being examined.
[0065] In the context of the present patent application, the term “centering position” means a position of the ocular inspection device 10 at which the following operating conditions occur: the optical axis a of the lens 2 of the ocular inspection device 10 passes through the center of the pupil EP, i.e., through the nominal position PN of the latter; and a front surface 20 of the lens 2 of the ocular inspection device 10 is located at a predefined optimal distance from the pupil EP, i.e., from the nominal position PN of the latter.
[0066] In general, the support base 11, the support structure 12 and the ocular inspection device 10 of the ophthalmic apparatus may be made according to known solutions and will be described below only with reference to the aspects of interest for the invention, for obvious reasons of descriptive brevity.
[0067] The centering system 3 of the ophthalmic apparatus comprises a detection unit 5 which includes at least one camera adapted to frame the pupil EP of the eye. The aforementioned detection unit 5 is adapted to provide a plurality of images Ii, E of the eye E taken from different angles with respect to the pupil EP at a generic moment of detection of the position PN of the pupil, during use of the ophthalmic apparatus. Preferably, each image Ii, I2 of the eye shows the pupil and a portion of the face surrounding the eye taken from a corresponding angle to the pupil (Figure 4).
[0068] Preferably, the detection unit 5 comprises a plurality of cameras 5A, 5B mounted on the ocular inspection device 10 and having optical axes b, c that are different from each other and oriented so that the pupil EP is framed by the aforementioned cameras, during use of the ophthalmic apparatus.
[0069] In principle, the optical axes b, c of the cameras may be oriented in any way (for example parallel to each other, skewed, or passing through the pupil) as needed, as long as the pupil is framed by the aforementioned cameras.
[0070] Preferably, the detection unit 5 comprises a pair of cameras 5A-5B to provide a pair of images Ii, I2, f the eye E taken from two different angles with respect to the pupil EP.
[0071] The cameras 5 A, 5B are integrally coupled to the ocular inspection device 10, advantageously at a front portion of the latter, facing the support structure 12.
[0072] The cameras 5 A, 5B therefore move together with the ocular inspection device 10.
[0073] Taking as a reference a normal operating position of the ophthalmic apparatus (illustrated in Figure 1), the cameras 5A, 5B are preferably located in a lower position with respect to the lens 2 so as to frame the patient's eye from below. Advantageously, they are arranged symmetrically, with respect to a vertical plane (not illustrated) passing through the optical axis a of the lens 2. According to other embodiments of the invention (not illustrated), the detection unit 5 comprises a single camera mounted on the ocular inspection device 10. The camera is movable (for example by virtue of a suitable electromechanical actuator) with respect to the ocular inspection device 10 (even though it moves in space together with the latter) so as to be able to assume optical axes different from one another. Such different optical axes are oriented so that the pupil EP is framed by the aforementioned camera according to different angles, during the use of the ophthalmic apparatus.
[0074] The detection unit 5 may also comprise an illuminator (not illustrated), which may advantageously be provided with one or more infrared light emitters adapted to emit light radiation towards the patient's eye E.
[0075] Such illuminator may be arranged to illuminate the external part of the eye E and partially the patient's face.
[0076] In general, the support detection unit 5 of the ophthalmic apparatus may be made according to known solutions and will be described below only with reference to the aspects of interest for the invention, for obvious reasons of descriptive brevity.
[0077] The centering system 3 of the ophthalmic apparatus comprises a data processing unit 6 capable of processing the images Ii, I2, provided by the detection unit 5 and generating, based on the acquired images, position data DSD indicative of a first relative position occupied by the pupil EP with respect to the ocular inspection device 10, at a certain detection instant.
[0078] By processing the images Ii, I2 provided by the detection unit 5, the data processing unit 6 is therefore able to provide, at a certain detection instant, the coordinates of a first position PA occupied by the pupil EP in a three-dimensional reference system (x’, y’, z’) integral with the ocular inspection device 10 (and therefore movable together with the cameras 5A, 5B).
[0079] The data processing unit 6 may comprise suitable microprocessor devices, FPGA circuits or other types of electronic circuits mounted on board a suitable electronic board. It is advantageously provided with appropriate computer resources to perform the intended functions. For example, as we will see better later, it may include in memory appropriate software application modules installed on board the microprocessor devices and executable by the latter to perform the required functions.
[0080] The centering system 3 comprises an actuation unit 4 adapted to receive and process the D3D position data D supplied by the data processing unit 6.
[0081] The actuation unit 4 is able to move the ocular inspection device 10 in space, along three axes in a fixed reference system (X, Y, Z), for example integral with the base 11 and the support structure 12 (Figure 1).
[0082] Based on the acquired position data D3D, the actuation unit 4 moves the inspection device 10 in space with respect to the nominal position PN of the pupil EP so that, after the movement of the aforementioned ocular inspection device, the pupil is in a second relative position PB with respect to the ocular inspection device 10 at which the ocular inspection device is centered with respect to the pupil EP (Figure 3).
[0083] In other words, with reference to a coordinate system (x’, y’, z’) integral with the ocular inspection device 10, the actuation unit 4 moves the ocular inspection device in space so that the patient's pupil EP moves from a first position PA, as detected by the detection unit 5, to a second predefined position PB corresponding to a centering position of the ocular inspection device (Figure 3).
[0084] For clarity, it is reiterated that the first position PA and the second position PB are relative positions of the pupil EP with respect to the ocular inspection device 10. In a coordinate system (x’, y’, z’) integral with the ocular inspection device, the first position PA is the position occupied by the pupil at a certain instant of detection while the second position PB is the position that the pupil EP should occupy with respect to the ocular inspection device 10 so that the latter is in the correct centering position. Obviously, the positions PA, PB of the pupil EP in a coordinate system (x’, y’, z’) integral with the ocular inspection device 10 are indicative of corresponding positions of the ocular inspection device with respect to the nominal position PN of the pupil, in a fixed reference system (X, Y, Z), for example integral with the base 11 and the support structure 12.
[0085] In principle, the second position PB of the pupil, in a coordinate system (x’, y’, z’) integral with the ocular inspection device 10, is predefined. However, it may change during the ocular examination due to movements of the patient's eye that cause corresponding changes in the nominal position PN of the pupil in a fixed (X, Y, Z) reference system.
[0086] The above-mentioned centering process is therefore repeated at subsequent instants of pupil position detection in which the first relative position PA of the pupil with respect to the ocular inspection device 10 is followed from time to time in order to position, in real time, the ocular inspection device in the correct centering position PB.
[0087] The actuation unit 4 preferably includes a control module 41 configured to process the acquired position data DSD and generate appropriate command signals to control a motorized mechanism 42 operatively connected to the ocular inspection device 10.
[0088] The control module 41 may comprise suitable microprocessor devices, FPGA circuits or other types of electronic circuits mounted on board a suitable electronic board. It is advantageously provided with appropriate computer resources to perform the expected functions and may be made according to known solutions. The motorized mechanism 42 may also be made according to known solutions. For example, it may comprise servomotors or stepper motors operatively connected to the ocular inspection device 10 by means of an appropriate kinematic chain.
[0089] In general, the actuation unit 4 of the ophthalmic apparatus may be made according to known solutions and will be described below only with reference to the aspects of interest for the invention, for obvious reasons of descriptive brevity.
[0090] According to the invention, the data processing unit 6 comprises an artificial intelligence (Al) module 61 configured to process the images Ii, E provided by the detection unit 5 and provide information M, D2D, Diiis indicative of the position occupied by the pupil EP.
[0091] Advantageously, the Al module 61 is configured to execute machine learning algorithms to analyze, at each detection instant during the operation of the ophthalmic apparatus, therefore during the ocular examination, the information sent by the detection unit 5 (i.e., the images Ii, I2) and provide real time information about the actual position of the pupil EP of the eye being examined.
[0092] The Al module 61 is capable of identifying the position of the pupil EP based on the features (visual patterns) of the images Ii, I2 not only when the pupil EP is fully shown in said images but also when the pupil EP is at least partially hidden in said images due to the presence of occlusions when the cameras 5 A, 5B of the detection unit framed the pupil.
[0093] The Al module 61 is further capable of predicting the position of the pupil EP based on visual patterns of the images Ii, E when the pupil EP is not shown in said images as the pupil was not included at least partially in the field-of-view of the cameras 5A, 5B of the detection unit. Advantageously, the Al module 61 is capable of using information (visual patterns of the image) derived from areas surrounding the pupil (e.g., iris, sclera, eyelashes, eyebrows, wrinkles) to calculate the position of the pupil EP when it analyses the images E, I2 provided by the detection unit 5.
[0094] In this way, the Al module 61 is capable of dealing with any obstacles (e.g. skin, eyebrows, eyelids) in the analyzed images Ii, I2, which partially or totally cover the pupil EP when the cameras 5A, 5B of the detection unit frame the pupil EP, and it is capable of calculating the position of the pupil EP accordingly.
[0095] Also, the Al module 61 is capable of recognizing situations in which the pupil EP is not shown in the analyzed images Ii, I2 because outside the field-of-view of the cameras 5A, 5B of the detection unit, and it is capable of calculating the position of the pupil EP also in these circumstances.
[0096] The information related to a position occupied by the pupil EP, which is provided by the Al module 61, can include one or more semantic segmentation maps of the eye.
[0097] Figure 5 illustrates a pair of semantic eye segmentation maps Mi, M2 corresponding to a pair of eye images taken from different angles. Each semantic segmentation map is essentially a binary image of the eye where the pupil is represented by pixels of one color and the rest of the image is represented by pixels of a different color. For example, in Figure 5, each semantic segmentation map Mi, M2 consists of a black and white image where the white pixels represent the pupil, and the black pixels represent everything else in the image.
[0098] The position information provided by the Al module 61 may be, for example:
[0099] M = [Mi, M2] wherein Mi, M2 are the semantic segmentation maps corresponding to the images Ii, I2 provided by the detection unit 5.
[0100] The information related to a position occupied by the pupil EP, which is provided by the Al module 61, can further include position data indicative of a position occupied by the pupil EP in a two-dimensional reference system, for example the reference system of the captured image. In this case, the position information provided by the Al module 61 may be, for example:
[0101] D2D = {(xi, yi), (x2, y2)} wherein (xi, yi), (x2, yi) are the pupil coordinates in the images Ii, I2 provided by the detection unit 5.
[0102] Optionally, the position information provided by the Al module 61 can also include a measurement value r indicative of the pupil radius. In this case, the position information D2D may be, for example, D2D = {(xi, yi), (X2, y2), r}.
[0103] The information related to a position occupied by the pupil EP, which is provided by the Al module 61, can further include position data indicative of a relative position of the pupil with respect to the ocular inspection device 10. In practice, the position information provided by the Al module 61 includes position data indicative of a position occupied by the pupil EP in a three- dimensional reference system (x’, y’, z’) integral with the ocular inspection device 10. In this case, the position information provided by the Al module 61 may be:
[0104] D3D = {(x, y, z)} wherein (x, y, z) are the coordinates of the pupil shown in the images Ii, I2 provided by the detection unit 5 in a three-dimensional reference system integral with the ocular inspection device 10 (figure 2B).
[0105] The nature of the information indicative of the position occupied by the pupil EP depends substantially on how the Al module 61 is structured and trained.
[0106] In the embodiments shown in the cited figures, the Al module 61 is configured to provide output information indicative of the position occupied by the pupil, which includes semantic segmentation maps of the eye or position data referred to a two-dimensional reference system. In this case, the data processing unit 6 preferably comprises a data processing module 62 adapted to process the information M, D2D provided by the artificial intelligence module 61 and provide position data D3D indicative of a first relative position PA occupied by the pupil with respect to the ocular inspection device 10, at a certain detection instant.
[0107] According to variant embodiments (not shown), however, the Al module is configured to directly provide position data D3D indicative of a relative position of the pupil with respect to the ocular inspection device 10. In this case, the Al module 61 can advantageously provide such information directly to the actuation unit 4, which processes it to move the ocular inspection device 10 in space as illustrated above.
[0108] Referring now to figures 6-12, an important aspect of the invention relates to a method 100 for creating an Al module 61 for processing images Ii, I2 of an eye E of a patient sent by a detection unit 5 of an ophthalmic apparatus. In operation, the Al module is capable of carrying out the above-described inferential tasks, i.e. , determining the position of the pupil EP of the eye under examination based on visual patterns of the images Ii, I2 sent by the detection unit 5. The steps of method 100, according to the invention, are described in the following with reference to figures 6-12. The steps of method 100, according to the invention, are now described referring to the most logical order shown in figure 11. However, in principle, they may be carried out with a different order with respect to the sequence shown in figure 11. According to the invention, the method 100 includes a step 101 of defining the architecture and the hyperparameters of the Al module 61 (figure 11).
[0109] In the definition step 101, the architecture desired for the Al module 61 is defined. This step, which is carried out before the training of the Al module begins, may include, for example, setting parameters such as the learning rate, the batch size, the type, and number of layers of the Al module 61.
[0110] In principle, the architecture of the Al module 61 may be of any type, according to the needs. According to preferred embodiments of the invention, however, such architecture may be a convolutional neural network (CNN), for example of the UNet type (figure 13). This choice is quite advantageous as such kind of architectures allows solving image segmentation tasks (i.e., given an input image, they provide an output classification pixel by pixel) unlike other types of convolutional neural networks solving classification problems.
[0111] Preferably, the computational architecture of the Al module 61 includes an encoder (or contraction path) processing an input image Ii, H of an eye E of a patient.
[0112] The encoder generates a sequence of feature maps related to the input image Ii, I2 through suitable convolutional layers and progressively reduces the size of these feature maps through suitable pooling layers. As an example, the encoder may include a sequence of consecutive blocks, each composed of two 3x3 convolutional layers followed by ReLu activations. Between a pair of consecutive blocks, a 2x2 pooling layer is provided, which halves the size of the generated feature maps by subsampling and discarding less informative features along the contraction path of the network. At each encoder block, the depth of the feature maps progressively increases while their spatial dimensions decrease, until each feature map is reduced to a size of 1x1 pixel.
[0113] Preferably, the architecture of the Al module 61 comprises a bottleneck in which the information (1x1 feature maps) obtained from the encoder is processed to obtain the requested information related to the position occupied by the pupil EP.
[0114] Preferably, the computational architecture of the Al module 61 comprises a fully connected network linked to the bottleneck. In this way, the coordinates D2D = (x, y) of the center of the pupil and, possibly, a measurement of the pupil radius r can be directly extracted.
[0115] When segmentation maps of the eye are provided by the Al module 61, the computational architecture of the Al module 61 advantageously comprises a decoder (or expansion path), which processes the information obtained by both the bottleneck and the contraction path to provide a segmentation map Mi, M2 of the eye as an output image.
[0116] The decoder generates a sequence of feature maps related to the output image Mi, M2 through suitable convolutional layers and progressively increases the spatial dimensions of these feature maps through suitable up-sampling layers, which are connected to the corresponding feature maps of the encoder via skip connections. The decoding process continues until the size foreseen for the output image Mi, M2 is reached. As an example, the decoder can include a sequence of consecutive blocks, each composed of a 2x2 up-sampling layer, which is concatenated with the above-mentioned skip connections, followed by two 3x3 convolutional layers with ReLu activations. At each block of the decoder, the depth of the neural network is progressively decreased until reaching depth=l in the output layer of the decoder.
[0117] The output layer of the neural network is preferably designed as a sigmoid activation layer, which provides pixel by pixel the probability of belonging to a class. Specifically, the sigmoid activation layer provides the probability of belonging to the pupil of the eye or the probability of belonging to parts of the eye different from the pupil.
[0118] Furthermore, in the architecture of the Al module 61, a batch normalization may be performed after each convolutional layer. Furthermore, before each pooling layer, a regularization dropout can be performed.
[0119] Furthermore, as an example, the standard convolutional blocks can be replaced with recurrent convolutional blocks to iteratively refine feature representations and potentially enhance extraction of contextual or spatially complex patterns.
[0120] In a practical implementation, if designed as explained above, the architecture of the Al module 61 can include about 950,000 parameters (949,000 of which are trainable).
[0121] According to the invention, the method 100 comprises a step 102 of preparation of a dataset Ds, used for defining one or more trainable parameters of the Al module 61. The dataset Ds advantageously includes a training dataset Dtrain, a validation dataset Dvaiand a test dataset Dtest, which include data to be used respectively for the training, the validation, and the testing of the Al module 61.
[0122] Referring now to figure 12, the dataset preparation step 102 preferably comprises a step 102a of acquiring a plurality of training images IIT, LT of an eye E of a patient by means of an ocular inspection device (figure 6).
[0123] In order to obtain a high variability of the training dataset Dtrain, the training images IIT, LT are conveniently captured in different observation setups. Advantageously, the acquired training images IIT, LT are captured for different patients, for different eyes (right, left) and for different physical conditions of the eyes.
[0124] The acquired training images IIT, LT can be framed in a manner similar to the images Ii, h provided by the detection unit during normal operation of the ophthalmic apparatus (Figure 4). In principle, however, they can be framed also in different ways.
[0125] Preferably, the acquired training images IIT, LT include images of a same eye taken from different angles with respect to the pupil. More preferably, the acquired images are framed at angles similar to those of the optical axes of the cameras 5A, 5B intended to frame the pupil during the ocular examination.
[0126] Preferably, the acquired training images IIT, I2T include images of a same eye with the pupil in various positions inside the eye socket and, consequently, with the pupil having various shapes (either circular or ellipsoidal) and showing / hiding different portions of the iris and the sclera. These images can be obtained by framing the patient’s eyes with the patients varying the direction of gaze.
[0127] Preferably, the acquired training images IIT, LT include images of a same eye in different light conditions, for example in presence of an artificial neon light, in presence of incandescent or LED bulbs behind or in front of the patient, or even in extreme natural light conditions, such in presence of a strong sunlight or in the dark.
[0128] Preferably, the above-mentioned dataset preparation step 102 includes the step 102e of selecting the training images IIT, I2T acquired in the acquisition step 102a.
[0129] The selection step 102e allows to purge the dataset Ds from duplicated or excessively similar images, which would not contribute with meaningful information to the training process Instead, it helps preventing certain configurations from excessively influencing the trainable parameters (weights) during the training process, at the expense of the less frequent configurations.
[0130] Advantageously, the selection step 102e can be carried out in a fully automated manner by executing known image processing algorithms designed to evaluate a similarity level between pairs of images and, possibly, discarding the images with an excessively high similarity level. Preferably, the above-mentioned dataset preparation step 102 comprises a step 102b of labelling the training images IIT, LT acquired in the acquisition step 102a (which can be possibly selected in step 102e to eliminate duplicated or similar images).
[0131] The labelling process 102b consists of associating control information (ground truth) CT to each acquired training image IIT, I2T. A plurality of data groups is formed during the labelling step 102b. Each data group comprises one or more training images IIT, LT of an eye and control information CT (label) associated with the image IIT, I2T of an eye.
[0132] According to some embodiments of the invention, the control information CT includes one or more eye segmentation maps. As explained above, each segmentation map is essentially a binary image of the eye where the pupil is represented by pixels of one color and the rest of the image is represented by pixels of a different color. For example, as shown in figure 5, a semantic segmentation map consists of a black and white image where the white pixels represent the pupil, and the black pixels represent everything else in the image.
[0133] According to other embodiments of the invention, the control information CT (label) includes position data indicative of a position occupied by the pupil in a two-dimensional reference system, for example the reference system of the captured image. The control information CT may also optionally include a measurement value r indicative of the pupil radius.
[0134] According to further embodiments of the invention, the control information CT includes position data indicative of a relative position of the pupil with respect to the ocular inspection device 10. In this case, the control information CT (label) includes position data indicative of a position occupied by the pupil EP in a three-dimensional reference system (x’, y’, z’) integral with the ocular inspection device 10.
[0135] Preferably, the labelling step 102b is carried out in a fully automated manner by executing suitable image processing algorithms, for example naive geometric algorithms, to identify the pupil in each acquired training image IIT, LT and provide control information CT to be associated to the processed image.
[0136] This solution is quite advantageous as it allows to process a huge amount of acquired training images and to create a huge number of distinct data groups (each including at least a training image IIT, I2T and the associated control information CT) in a reasonable time.
[0137] Preferably, the above-mentioned dataset preparation step 102 comprises a step 102c of checking the control information CT included in the data groups obtained during the above-mentioned labelling step 102b.
[0138] The checking step 102c is preferably carried manually with the aid of suitable support software. For each data group, human annotators examine the control information (label) associated to each corresponding acquired training image IIT, T, which was proposed by the geometric algorithms during the labelling process (step 102b).
[0139] The human annotators may discard training images with completely incorrect annotations, accept accurate labels, or manually correct partially incorrect labels.
[0140] For example, when the control information CT includes a segmentation map, the human annotators can identify the ellipsoid representing the pupil in each acquired training image IIT, I2T and correctly draw the corresponding blob in the corresponding segmentation map representing the ground truth. The human annotators can also edit the coordinates (x, y) of the center of the pupil.
[0141] Both these pieces of information are part of the ground truth information used by the Al module 61 to perform a semantic segmentation process (if the information to be provided includes a segmentation map Mi, M2 of the eye) or a regression process (if the information to be provided includes the coordinates of the center of the pupil).
[0142] The above-described checking step 102c is particularly advantageous as it allows to remarkably increase the quality of the data included in the dataset Ds used to define the trainable parameters of the Al module 61. The automated labelling process 102b may, in fact, be affected by a relatively high percentage of errors, which would negatively affect the performance of the Al module 61. Improving the quality of the ground truth information used to train the Al module 61 remarkably improves the final performance of the latter, once trained.
[0143] Preferably, the above-mentioned dataset preparation step 102 comprises a step 102d of partitioning the data groups obtained in the checking step 102c to obtain the above-mentioned training dataset Dtrain, validation dataset Dvaiand test dataset Dtest. Advantageously, each dataset includes images of eyes of different patients. This allows preventing that the same information used to train the Al module 61 is used also for validation and testing (avoid data-leakage).
[0144] EXAMPLE
[0145] An example of practical implementation of the dataset preparation process 102 to obtain a dataset Ds, which can be used to define the trainable parameters of the Al module 61, is proposed in the following.
[0146] During the above-mentioned acquisition step 102a, 28 healthy subjects, aged between 20 and 55 years, were considered and examined to acquire images of their eyes. 14 subjects were examined for the right eye while further 14 subjects were examined for the left eye. In addition, the eyes of 4 subjects were pharmacologically dilated by administering ophthalmic drops. For each subject, between 3000 and 6000 images of the eye were finally acquired.
[0147] At the end of the filtering process 102e, labelling process 102b and checking process 102c, 70000 training images were included in the dataset Ds. Each training image was included in a data group in association with its corresponding control information (ground truth).
[0148] The data groups were partitioned into three distinct sets (training, validation, and test sets), each including images of distinct subjects. Overall, 50000 images were included in the training dataset Dtrain, 10000 images in the validation dataset Dvaiand 10000 images in the test dataset Dtest. The above-mentioned datasets were then used to train, validate and test the Al module 61 . According to the invention, method 100 comprises a step 103 of carrying out one or more training cycles (i.e., epochs) of the Al module 61 by using the training dataset Dtrain.
[0149] Conveniently, the training process 103 of the Al module 61 includes a plurality of training cycles (epochs). At each training cycle, the entire training dataset Dtrain is evaluated in subsets of training data (i.e., batches).
[0150] According to the invention, method 100 comprises a step 104 of data augmentation, which consists in modifying one or more training images IIT, LT of one or more batches of the training dataset Dtrain.
[0151] Preferably, the modification (augmentation) of the training images is carried out randomly, according to predefined modification parameters, on the batches of the training dataset Dtrain at one or more training cycles (preferably each training cycle) of the training process 103.
[0152] Step 104 of the method 100 allows generating a virtually infinite number of variants of the training dataset Dtrain. Advantageously, the training images IIT, LT of the training dataset Dtrain are modified through appropriate data augmentation techniques, which may include, for example, random horizontal flip, rotation, brightness, and contrast adjustment.
[0153] The training images IIT, LT can be suitably modified to simulate the presence of artifacts that degrade the quality of the eye images. This solution forces the Al module 61 to learn to recognize the pupil in an eye image even in the presence of artifacts that may change the appearance of the image. Such artifacts may include, for example, make-up on the patient's face, spurious light sources, shadows, moles, skin blemishes, tattoos, and so on.
[0154] An important aspect of the invention is that, at step 104 of method 100, one or more training images IIT, LT of the training dataset Dtrain are modified in such a way to simulate one or more among the following circumstances: one or more circumstances, in which the pupil EP was at least partially hidden by occlusions, when a camera framed the eye of a patient; one or more circumstances, in which the pupil EP was not included at least partially in the field-of-view of an image capturing the eye of a patient.
[0155] Preferably, the modified training images simulating the presence of occlusions can be obtained through the application of masks (with various shapes and sizes) on the acquired training images. Said masks can partially or completely hide the pupil EP of the eye.
[0156] Figures 7-10, 10A show some examples of training images IIT, ET modified to simulate the presence of a pupil occlusion. As it is evident from the above-mentioned figures, the modified training images may be advantageously produced in such a way as to simulate the presence of occlusions of different types and with different shapes, size, and color, which may partially or completely hide the pupil.
[0157] Preferably, the modified training images simulating a partial or a complete failure to frame the pupil EP can be obtained through suitable image translation techniques (applied to both the images and the ground truth) that allow moving the pupil outside the field of view of the acquired training images.
[0158] It is worth noting that the augmentation of training images to simulate the presence of occlusions and / or a failure to frame the pupil would not be possible during the building up step 102 as it would be virtually impossible to define a ground truth for an image showing a hidden or missing pupil.
[0159] According to the solution proposed by the invention, instead, the above-described augmentation techniques are applied on training images, which have been already associated with the correct control information CT (label).
[0160] A training process carried out using these modified training images forces the Al module 61 to learn to recognize the pupil using the information (features) available from the context surrounding the pupil (iris, sclera, eyelashes, eyebrows, wrinkles, and so on). In this way, once trained, the Al module 61 will be capable of identifying the position of the pupil regardless if it is fully shown or partially hidden in the images acquired during an eye examination session, and will be capable of predicting the position of the pupil when this latter is not shown at all in said images since it is not included in the field-of-view of said images.
[0161] According to the invention, the method 100 comprises a step 105 of validation of the Al module 61 at each training cycle using the data included in the validation dataset Dvai.
[0162] For instance, this validation step may involve generating predictions on the validation dataset and computing a corresponding loss function. By comparing both the losses on the training dataset and on the validation dataset, it is possible to assess whether the model is starting to overfit.
[0163] This comparison can also be used to implement an early stopping mechanism, where training is halted before reaching the maximum number of epochs if the validation loss stops improving or starts worsening, even if the training loss is still improving. Optionally, this mechanism can include a “patience counter”, which allows training to continue for a few additional epochs in case the model improves again after a short stagnation in validation performance. According to the invention, the method 100 comprises a step 106 of testing the Al module 61 using the data included in the test dataset Dtest.
[0164] The testing process of the Al module can be performed in a standard manner. It consists of evaluating the final trained model on completely unseen data (i.e., the test dataset Dtest, which - unlike the training and validation datasets Dtest, Dvai- has been never used during any phase of the model training) to objectively assess the model’s generalization ability.
[0165] After the test process 106, the (trained) Al module 61 is ready for operation.
[0166] The general operation of the centering system 3 of the ophthalmic apparatus 1, according to the embodiments illustrated in the cited figures, is described below.
[0167] At an initial detection instant (e.g., at the beginning of the ocular examination), cameras 5A, 5B of the detection unit 5 provide a pair of images Ii, I2 of the eye of the patient being examined. The data processing unit 6 processes the images provided by the detection unit 5 and outputs position data DID indicative of a first relative position PA occupied by the pupil EP with respect to the ocular inspection device 10.
[0168] The actuation unit 4 receives and processes the position data DSD provided by the data processing unit 6 and moves the ocular inspection device 10 in space so that, after the movement of the aforementioned ocular inspection device, the pupil EP is in a second relative position PB with respect to the ocular inspection device 10, at which point the ocular inspection device 10 is centered with the pupil EP.
[0169] The above-mentioned centering process is repeated at subsequent times of pupil position detection. Whenever the first relative position PA of the pupil with respect to the ocular inspection device 10, as detected, does not match with the second relative position PB corresponding to the correct centering position of the ocular inspection device 10 (for example due to movements of the patient's eye), the actuation unit 4 moves the ocular inspection device 10 so as to minimize (in practice cancel) the distance between the relative positions of the pupil PA, PB mentioned above. In this way, after the movement by means of the actuation means 4, the ocular inspection device 10 is in the correct centering position with the pupil of the eye of the patient being examined.
[0170] It has been seen in practice that the ophthalmic apparatus 1, according to the invention, allows the described drawbacks of the prior art to be solved, achieving the intended objects.
[0171] The ophthalmic apparatus 1 includes improved centering means 3 which allow the position of the pupil to be assessed, according to three reference axes, with high measurement precision and continuously, even during the ocular examination.
[0172] By using processing means provided with an artificial intelligence module 61, the centering means 3 of the apparatus allow the ocular inspection device 10 to be aligned with the pupil of the eye being examined with a high level of accuracy. Training using augmented training images modified using data augmentation techniques allows the artificial intelligence module to calculate the position of the pupil of the eye being examined, based on the images sent by the detection unit 5, with levels of precision much higher than those obtainable with traditional image processing methodologies, in particular when the eyelids interfere with the imaging of the pupil or when the imaging of the patient's eye by the camera is carried out in particular conditions, for example in the presence of make-up on the patient's face, spurious light sources, shadow areas, partial occlusions of the eye, moles, skin blemishes, tattoos, and so on.
[0173] The ophthalmic device, according to the invention, is therefore characterized by a high and constant operational reliability (which may be advantageously improved over time thanks to the training of the artificial intelligence module) and does not require complex calibration and set-up procedures for its installation.
[0174] The ophthalmic apparatus, according to the invention, has a relatively simple structure and may be manufactured industrially at competitive costs.
Claims
CLAIMS1. Ophthalmic apparatus (1) comprising an ocular inspection device (10) and a centering system (3) capable of positioning said ocular inspection device in a centering position with respect to the pupil (EP) of an eye (E) of a patient, said centering system (3) including: a detection unit (5) comprising at least a camera (5A, 5B) adapted to frame the pupil (EP), said detection unit being adapted to provide a plurality of images (Ii, I2) of the eye (E) taken from different angles with respect to the pupil (EP) at a certain detection instant; a data processing unit (6) capable of acquiring the images (Ii, I2 ) provided by said detection unit and provide, based on the acquired images, position data (DSD) indicative of a first relative position (PA) occupied by the pupil (EP) with respect to said ocular inspection device (10) at said detection instant; an actuation unit (4) adapted to process the position data (DID) provided by said data processing unit and move, based on the acquired position data, said ocular inspection device (10) in space so that, after the movement of said ocular inspection device, the pupil (EP) is in a second relative position (PB) with respect to said ocular inspection device, at which said ocular inspection device is centered on the pupil (EP); characterized in that said data processing unit (6) includes an artificial intelligence module (61) configured to process the images (Ii, I2) provided by said detection unit (5) and provide information (M, D2D, DSD) indicative of a position occupied by the pupil (EP), wherein, in operation, said artificial intelligence module (61) is capable of identifying the position of the pupil (EP) based on visual patterns present in said images (Ii, I2) both when the pupil (EP) is fully shown in said images (Ii, I2) and when the pupil (EP) is at least partially hidden in said images, wherein, in operation, said artificial intelligence module (61) is capable of predicting the position of the pupil (EP) based on visual patterns present in said images (Ii, I2) when the pupil (EP) is not included at least partially in the field-of-view of said at least a camera.
2. Apparatus, according to claim 1, characterized in that the information related to a position occupied by the pupil (EP), provided by said artificial intelligence module (61), includes one or more among:- one or more semantic segmentation maps (Mi, M2) of the eye;- position data (D2D) indicative of a position occupied by the pupil in a two- dimensional reference system;- position data (Dsn) indicative of a relative position (PA) occupied by the pupil with respect to said ocular inspection device (10).
3. Apparatus, according to claim 2, characterized in that said data processing unit (6) includes a data processing module (62) adapted to process information (M, DZD) indicative of a position occupied by the pupil (EP), provided by said module artificial intelligence (61), and to provide said position data (DSD) indicative of a relative position (PA) occupied by the pupil with respect to said ocular inspection device (10).
4. Apparatus, according to one of the previous claims, characterized in that said detection unit (5) comprises a plurality of cameras (5A, 5B) mounted on said ocular inspection device (10) and having optical axes (b, c) different one from another and oriented in such a way that the pupil (EP) can be framed by said cameras.
5. Apparatus, according to one of the claims from 1 to 3, characterized in that said detection unit (5) comprises a single camera mounted on said ocular inspection device (10) and movable with respect to said ocular inspection device.
6. Apparatus, according to one of the previous claims, characterized in that said ocular inspection device (10) is a device for inspection of the ocular fundus.
7. Apparatus, according to claim 6, characterized in that said ocular inspection device (10) is a fundus camera, a FAP device, or an OCT inspection device.
8. A method (100) for providing an artificial intelligent module (61) for processing the images (Ii, h) of an eye provided by a detection unit (5) of an ophthalmic apparatus, wherein, in operation, said artificial intelligence module (61) provides information (M, D2D, DSD) indicative of a position occupied by the pupil (EP), wherein, in operation, said artificial intelligence module (61) is capable of identifying the position of the pupil (EP) based on visual patterns present in said images (Ii, I2) both when the pupil (EP) is fully shown in said images (Ii, I2) and when the pupil (EP) is at least partially hidden in said images, wherein, in operation, said artificial intelligence module (61) is capable of predicting the position of the pupil (EP) based on visual patterns present in said images (Ii, I2) when the pupil (EP) is not included at least partially in the field-of-view of a camera capturing the eye of a patient, wherein said method (100) comprises the following steps:- defining (101) architecture and hyperparameters of said artificial intelligence module(61);- preparing (102) a dataset (Ds) to learn the trainable parameters of said artificial intelligence module (61), wherein said dataset includes a training dataset (Dtrain), a validation dataset (Dvai) and a test dataset (Dtest);- carrying out one or more training cycles (103) of said artificial intelligence module (61) by using said training dataset (Dtrain);- modifying (104), at one or more said training cycles, one or more training images (IIT, LT) of said training dataset (Dtrain), wherein the modified training images simulate at least one among:- one or more circumstances, in which the pupil (EP) is at least partially hidden by occlusions; and- one or more circumstances, in which the pupil (EP) is not included at least partially in the field of a camera capturing the eye of a patient;- validating (105) said artificial intelligence module (61) by using said validation dataset (Dvai) during the training of such artificial intelligence module.;- testing (106) said artificial intelligence module (61) by using said testing dataset (Dtest), once said artificial intelligence module has been trained.
9. Method, according to claim 8, characterized in that said dataset preparation step (102) comprises:- acquiring (102a) training images (IIT, ET) of an eye (E) of a patient by means of an ocular inspection device, wherein said training images are acquired in different observation setups;- labelling (102b) the acquired training images (IIT, ET) and forming a plurality of data groups, each data group comprising at least a training image (IIT, ET) and control information (CT) associated with said at least a training image (IIT, ET);- checking (102c) the control information (CT) included in said data groups;- partitioning (102d) the data groups to obtain said training dataset (Dtrain), said validation dataset (Dvai) and said test dataset (Dtest).
10. Method, according to claim 9, characterized in that said dataset preparation step (102) comprises selecting (102e) the acquired training images (IIT, LT).
11. Method, according to one of the claims from 8 to 10, characterized in that said control information (CT) includes one or more among:- one or more semantic segmentation maps (Mi, M2) of the eye;- position data (D2D) indicative of a position occupied by the pupil in a two-dimensional reference system;- position data (DSD) indicative of a relative position (PA) occupied by the pupil with respect to said ocular inspection device (10).
Citation Information
Patent Citations
Ophthalmologic apparatus
EP1442698A1
Fundus camera
US20050117115A1
Alignment guidance user interface system
WO2022117646A1
Ophthalmologic apparatus
US20240398223A1
System and method for determining pupil center based on convolutional neural networks
US20250005908A1