Retinal image capture

JP7904899B2Active Publication Date: 2026-08-13MEDIOS TECHNOLOGIES PTE LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-10-05
Publication Date
2026-08-13

Smart Images

  • Figure 0007904899000001
    Figure 0007904899000001
  • Figure 0007904899000002
    Figure 0007904899000002
  • Figure 0007904899000003
    Figure 0007904899000003
Patent Text Reader

Abstract

An approach for automated retinal image capturing is described. In one example, a plurality of target frames recorded by an imaging device are received, the plurality of target frames relating to a subject's retina. The plurality of target frames are then analyzed based on a trained analysis model. For example, the analysis model is trained based on a set of training images, the analysis model incorporating a set of confidence weights based on visual attributes and annotations of the set of training images. Based on the analysis of the plurality of target frames, an imaging device can be triggered to capture a retinal image of the subject's retina.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Relates to retinal image capturing.

Background Art

[0002] An image capturing device may be used to capture a retinal image for a medical examination of an eye. Capturing a retinal image in a handheld mode using an image capturing device can be a cumbersome process. In such cases, the image capturing device may have to be correctly positioned at the correct distance from the eye being targeted. Moreover, the image capturing device may have to be precisely pointed through the pupil of the eye in order to capture the retinal image. This requires the user of the image capturing device to evaluate the ambient luminance and accordingly adjust the image capturing device to capture the retinal image.

Summary of the Invention

Means for Solving the Problems

[0003] In the following detailed description, reference is made to the drawings.

Brief Description of the Drawings

[0004] [Figure 1] A diagram showing an exemplary system for automated retinal image capturing based on an analysis model, according to an example of the present subject matter. [Figure 2] A diagram showing a training system for training an analysis model, according to an example of the present subject matter. [Figure 3] A diagram showing a set of training images for training an analysis model, according to an example of the present subject matter. [Figure 4] A diagram showing a set of training images for training an analysis model, according to an example of the present subject matter. [Figure 5]This figure shows an evaluation system that implements an analytical model, as an example of the subject matter. [Figure 6] This figure illustrates an exemplary method for training an analytical model in a training system, using an example from this subject. [Figure 7] This figure shows a system environment that implements a non-temporary computer-readable medium for automated retinal image capture based on an analysis model, as an example of the subject of this paper. [Modes for carrying out the invention]

[0005] Throughout the drawings, the same reference number designates similar, but not necessarily identical, elements. The drawings provide examples and / or implementations that correspond to the description, but the description is not limited to the examples and / or implementations given in the drawings.

[0006] An imaging device can be a device capable of capturing and storing still or moving images. In some cases, an imaging device can also perform image processing to output images. An imaging device may include multiple components for recording, storing, manipulating, viewing, and transmitting visual images. Examples of components of an imaging device may include, but are not limited to, optical apertures, imaging lenses, microlens arrays, imaging elements, processors, and controllers. Examples of imaging devices include, but are not limited to, still cameras, camcorders, video cameras, and 3D cameras.

[0007] Imaging devices can be used for a variety of medical applications. They can be used to image various parts of the human body, such as the eye of a subject. For example, imaging devices can be used for retinal imaging for diagnostic and therapeutic purposes. For this purpose, a retinal image refers to a digital image of the back of the eye of a subject. The retinal image may capture or show the retina, optic disc, and blood vessels inside the eye of the subject. Such retinal imaging can be used for eye examinations or the diagnosis of eye diseases.

[0008] Typically, imaging devices are handheld to capture a retinal image of the subject's eye. Capturing a retinal image using a handheld imaging device can require considerable attention. In particular, users of imaging devices may need to be trained to be able to capture a retinal image of the eye for eye examination. Furthermore, users may need to position the imaging device at the correct working distance and point it through the center of the pupil of the eye. Users may need to evaluate and adjust the imaging device to capture a precise retinal image based, for example, on the working distance and the ambient lighting or brightness. The user should hold the imaging device steadily at the desired correct working distance to capture a retinal image. Any deviation in accommodation or any movement of the imaging device can result in an inaccurate or unusable retinal image. Additionally, the difficulty in capturing a retinal image increases if the subject's pupil is small or the subject is uncooperative.

[0009] In some cases, a retinal image can be captured by directing near-infrared light into the retina. The user of the imaging device can then inspect the device's field of view and trigger the capture when the correct working distance and positioning are achieved. This reduces the need to dilate the subject's pupil. However, the user may still need to adjust the ambient brightness or illumination to capture an accurate retinal image. Furthermore, the user may still need to evaluate the brightness distribution across the imaging device's field of view based on their experience. Such criteria may differ among technicians, and therefore, the quality of the images thus captured may vary.

[0010] Retinal image capture can be simplified by triggering the capture of retinal images under correct operating conditions through an automated mechanism. In this regard, retinal images captured under correct operating conditions can be precise and have high resolution and brightness, making them suitable for medical applications. Conventional automated capture features of imaging devices may rely on optics and the detection of the trajectory of light rays emanating from the retina to capture retinal images. In another example, automated capture features may rely on image processing techniques to detect the presence of retinal features within the field of view of the imaging device. However, the operation and function of such automated capture features may not accurately or clearly represent the retina. In particular, such automated capture features cannot guarantee the correct illumination of the retina of the eye. As a result, it may be difficult for the examiner of the retinal image to draw accurate conclusions from the retinal image, thereby rendering the automated capture feature for retinal imaging useless. Moreover, such automated capture features may require additional processing devices, thereby making the imaging device bulky and complex. This can hinder the experience of the user and the subject using the imaging device.

[0011] This paper describes a method for automated retinal image capturing based on an analytical model. In this context, the analytical model can be used to trigger the capture of retinal images based on visual attributes within the field of view of an imaging device. The analytical model utilizes machine learning techniques to provide a mechanism by which the visual attributes within the field of view of the imaging device can be monitored. In such cases, the analytical model is first trained to evaluate the visual attributes within the field of view of the imaging device, and therefore can trigger automated retinal image capturing based on that evaluation. Examples of such visual attributes include, but are not limited to, light distribution parameters, light intensity, brightness, contrast, sharpness, and shape.

[0012] In one example, an imaging device may record multiple target frames of the target retina. In particular, the imaging device may be positioned in the path of light reflected from the retina in order to record multiple target frames. As can be understood, the imaging device may collect the light rays reflected from the retina and redirect them to a single point, namely the focal point of the imaging device. In one example, the retina of the eye may be illuminated with near-field infrared light.

[0013] For this purpose, the imaging device may record multiple target frames in the real-time mode of the imaging device. These multiple target frames may be recorded before the actual capture of the retinal image. For example, multiple target frames may be recorded by the imaging device while it is adjusting to the target eye or retina.

[0014] During operation, the system may receive multiple target frames from an imaging device. The system may analyze multiple target frames based on an analysis model. The analysis model may first be trained on a set of training images. In one example, the analysis model is trained on the visual attributes and annotations corresponding to the set of training images. In this regard, the analysis model can be considered to incorporate a set of confidence weights based on the visual attributes and annotations for each of the set of training images. In one example, the confidence weights for the training images may correspond to and be determined based on the correlation between the visual attributes and annotations for the training images.

[0015] In one example, the training image set could correspond to infrared images of the retina captured by an imaging device in real time. For instance, training an analytical model might involve using a large set of training images by a training system on which the analytical model can be trained. Furthermore, the analytical model might be trained once before deploying the trained analytical model on the system.

[0016] Once trained, the set of confidence weights and the trained analysis model can be deployed on the system. The analysis model can then be used to analyze multiple target frames recorded by the imaging device. For example, the analysis of multiple target frames is performed in real time, i.e., as the target frames are registered. Based on the set of confidence weights and the multiple target frames, the analysis model on the system only needs to determine whether any of the multiple target frames corresponds to a correctly positioned image in order to capture a retinal image. In one example, for a correctly positioned image, the imaging device may be positioned at the correct working distance and properly illuminated within the imaging device's field of view.

[0017] When the imaging device determines that the target frame corresponds to a correctly positioned image, it may be triggered to capture a retinal image of the target retina. Such retinal image capturing is performed under the correct operating conditions of the imaging device. For example, the correct operating conditions may be achieved when the imaging device is held at the correct working distance, the imaging device's field of view is properly illuminated, and the visual attributes of the retinal image are correct.

[0018] The automated retinal image capturing system described in this subject utilizes a machine learning-based analytical model to analyze target frames for capturing retinal images. The retinal images are therefore captured under correct operating conditions by the imaging device. The resulting retinal images can be precise and high-resolution, making them effective for medical applications. In this way, any noise or unwanted artifacts are substantially reduced in the automatically captured retinal images. In some cases, the technique for automated retinal image capturing, i.e., the machine learning model, can be deployed on handheld systems such as smartphones and digital cameras. Thus, precise retinal images can be captured using existing imaging techniques. Additionally, the use of the machine learning model does not complicate the system; rather, it improves the system's operability or ease of use. Therefore, precise retinal images can be automatically captured without any substantial increase in cost.

[0019] Further description of this subject matter is provided with reference to the attached drawings. Wherever possible, the same reference numerals are used in the drawings and the following description to refer to the same or similar parts. Note that the descriptions and drawings are merely illustrative of the principles of this subject matter. It will be understood that various arrangements encompassing the principles of this subject matter may be devised, even though they are not explicitly described or shown herein. Moreover, all statements herein describing the principles, embodiments, and examples of this subject matter, as well as specific examples thereof, are intended to encompass their equivalents.

[0020] The exemplary systems will be implemented in detail with reference to Figures 1 to 7. While the described systems may be implemented in any number of different electronic devices, environments, and / or implementations, examples are provided within the context of the exemplary devices described below. Note that the drawings of this subject matter shown herein are illustrative and should not be construed as limiting the scope of the claimed subject matter.

[0021] FIG. 1 shows an exemplary system 102 for automated retinal image capturing based on an analysis model, according to an example of the present subject matter. System 102 includes a processor 104 and a machine-readable storage medium 106 coupled to and thereby accessible by the processor 104. System 102 may be a computing system such as a memory array, a server, a desktop computer, a laptop computer, a smartphone, a distributed computing system, etc. Although not shown, system 102 may include other components such as an interface for communicating via a network or with an external memory or computing device, display, input / output interface, operating system, application, data, etc., which are not described for the sake of brevity.

[0022] Processor 104 may be implemented as a dedicated processor, a shared processor, or a plurality of individual processors, some of which may be shared. Machine-readable storage medium 106 may be communicatively coupled to processor 104. Among other capabilities, processor 104 may fetch and execute computer-readable instructions including instructions 108 stored in machine-readable storage medium 106. Machine-readable storage medium 106 may include a non-transitory computer-readable medium including, for example, volatile memory such as RAM or non-volatile memory such as EPROM, flash memory, etc. Instructions 108 may be executed to determine the occurrence of an anomaly in a target computing device based on an analysis of the current operating parameters of the target computing device.

[0023] In one example, the processor 104 may fetch and execute the instruction 108. For example, as a result of the execution of the instruction 110, the system 102 may receive a plurality of target frames from the imaging device. The plurality of target frames may relate to the retina of the subject or patient. The plurality of target frames may be captured in the real-time mode of the imaging device. In particular, the plurality of target frames may correspond to image frames of the retina of the subject from different positions and angles of the imaging device. The plurality of target frames may correspond to image frames within the field of view of the imaging device captured in real time. In one example, such a plurality of target frames may be temporarily stored in a memory such as the machine-readable storage medium 106 of the system 102.

[0024] The plurality of target frames of the retina may be analyzed based on an analysis model as a result of the execution of the instruction 112. The analysis model may be a machine learning model. As described above, the analysis model may be trained prior to being used to analyze the plurality of target frames. In particular, the analysis model may be trained using a set of training images. In one example, the analysis model may be trained prior to use in a training system, which may be different from the system 102.

[0025] Once trained, system 102 can use an analysis model to analyze multiple target frames of the retina received from the imaging device in order to determine whether the imaging device is properly configured to capture retinal images. In particular, system 102 can analyze multiple target frames in real time based on the trained analysis model. In this regard, system 102 only needs to analyze the visual attributes corresponding to the target frames based on or by using the analysis model. Based on the visual attributes of the target frames, system 102 can determine whether the target frames correspond to correctly positioned images in order to determine whether the imaging device is properly configured to capture retinal images.

[0026] Once multiple target frames are analyzed in real time, system 102 only needs to determine whether at least one of the target frames corresponds to a correctly positioned image. Subsequently, command 114, when executed by system 102, may cause the imaging device to capture a retinal image of the target retina. The retinal image is captured immediately upon registration of the target frame corresponding to the correctly positioned image. Note that the correctly positioned image may correspond to an image that properly captures the retina, for example, from the correct working distance between the imaging device and the retina, while the field of view of the imaging device is properly illuminated, i.e., when no unwanted artifacts are present. The retinal image may be captured based on the analysis of multiple target frames. This method is just one of many other examples that can be used for automated retinal image capturing. Such other methods may be used without limiting the scope of this subject.

[0027] The techniques described above, implemented as a result of the execution of instruction 108, may be carried out by different programmable entities. Such programmable entities may be implemented through a computing system that can be implemented on either a standalone computing device or on multiple computing devices. As described, various examples of this subject are described in the context of a computing system for training a neural network-based model and then using the neural network model for automated retinal image capture based on the analysis of multiple target frames. For example, such a neural network-based model is an analytical model that is trained and deployed on a computing system for automated retinal image capture. These and other examples are described further in relation to other figures.

[0028] Figure 2 shows a training system 202 for training an analysis model 204, as an example of this subject. The training system 202 includes a processor and memory (not shown) for training the analysis model 204. In some examples, the training system may be a processor-based system. Furthermore, the training system 202 (referred to as system 202) may communicate with a training data repository 208 through a network 206.

[0029] Network 206 may be a private or public network and may be implemented as a wired network, a wireless network, or a combination of wired and wireless networks. Network 206 may also include a collection of individual networks that are interconnected with each other and function as a single large-scale network, such as the Internet. Examples of such individual networks may include, but are not limited to, Global System for Mobile Communications (GSM) networks, Universal Mobile Communications System (UMTS) networks, Personal Communications Services (PCS) networks, Time Division Multiple Access (TDMA) networks, Code Division Multiple Access (CDMA) networks, Next Generation Networks (NGNs), Public Switched Telephone Networks (PSTNs), Long-Term Evolution (LTE), and Integrated Services Digital Networks (ISDNs).

[0030] The training data repository 208 (referred to as repository 208) may be a machine-readable storage medium. Repository 208 may be connected to and accessible by system 202. Repository 208 may include non-temporary computer-readable media, such as volatile memory like RAM, or non-volatile memory like EPROM or flash memory. Repository 208 may store training data, specifically a set of training images 210 and annotations 212. System 202 may use the set of training images 210 and annotations 212 to train the analysis model 204.

[0031] Furthermore, system 202 may include a training engine 214. The training engine 214 (referred to as engine 214) or any other engine in system 202 may be implemented as a combination of hardware and programming, for example, as programmable instructions for implementing various functionalities. In the examples described herein, such a combination of hardware and programming may be implemented in several different ways. For example, the programming for engine 214 may be executable instructions, such as instruction 216. Such instructions may be stored in a non-temporary machine-readable storage medium that can be coupled to system 202 either directly or indirectly (for example, through networked means). In one example, engine 214 may include processing resources, for example, a single processor or a combination of multiple processors, for executing such instructions. In this example, the non-temporary machine-readable storage medium may store instructions, such as instruction 216, and these instructions, when executed by the processing resources, may implement engine 214. In another example, the training engine 214 may be implemented as an electronic circuit mechanism.

[0032] Data 218 may include a set of training images 210 and annotations 212 received from repository 208 by system 202. Data 218 may further include a set of confidence weights 220 and other data 222. Each of the training images 210 may have a corresponding visual attribute and other parameters on which the analysis model 204 can be trained. System 202 may further include instructions 216 for training the analysis model 204 based on the visual attributes and other parameters of the training image set 210. The visual attributes and other parameters may include data or values ​​of different attributes relating to the training image set 210. The visual attributes and other parameters may be derived by processing the training image set 210 as a result of executing instructions 216, or by an engine such as the training engine 214.

[0033] During operation, system 202 may receive a set of training images 210. As can be understood, each of the training images 210 may have attributes such as visual attributes and other parameters. In one example, system 202 may analyze the set of training images 210 to determine the corresponding attributes.

[0034] Visual attributes for training images may correspond to patterns of image segments for the training images that describe several characteristic properties. In some examples, such visual attributes may be any combination of the appearance, shape, or layout of segments within the training image. For example, system 202 may process each of the set of training images 210 to identify the corresponding visual attributes. As can be understood, system 202 may perform such processing of the set of training images 210 either by itself or implicitly.

[0035] In accordance with this subject, the visual attributes of the training images from the training image set 210 may correspond to the visual attributes across the retinal boundary of the eye, the visual attributes at the center of the retina, the visual attributes of the pupil of the eye, the visual attributes of the optic disc, and the visual attributes of the blood vessels within the eye. Examples of visual attributes of the training image set 210 may include, but are not limited to, sharpness, resolution, brightness, light distribution parameters, light intensity, contrast, shape, size, color, texture, highlights, saturation, structure, and shadows. The visual attributes corresponding to the training images may relate to the retina, optic disc, pupil, and blood vessels of the eye captured within the training images.

[0036] It should be noted that the manner in which visual attributes occur within training images may correspond to the operating conditions under which the training images under consideration were captured. When the operating conditions are changed, several changes in visual attributes may occur and may be present in the training images.

[0037] Each of the training image set 210 may be associated with a corresponding annotation stored as annotation 212 on which the analysis model 204 can be trained. In this regard, annotation 212 may indicate from the training image set 210 that a corresponding image corresponds to either a correctly positioned image or an incorrectly positioned image. A correctly positioned image may be an intermediary image frame that corresponds to a correct position suitable for capture. On the other hand, an incorrectly positioned image may be an intermediary image frame that corresponds to an incorrect position that is neither suitable for capture nor use.

[0038] For this purpose, annotations 212 may be assigned to the set of training images 210 to indicate that the set of training images 210 corresponds to either a correctly positioned image that fits the capture or an incorrectly positioned image that does not fit the capture. In one example, annotations 212 may be labels, where correctly positioned images from the set of training images 210 may be labeled "correct," and incorrectly positioned images from the set of training images 210 may be labeled "incorrect."

[0039] In one example, footnote 212 may also specify a working distance that can be assigned to a set of training images 210. In particular, the working distance for the training images may correspond to the distance between the imaging device and the eye from which the image is captured. For example, such a distance may be expressed as a numerical value in micrometers, millimeters, centimeters, meters, etc. In another example, footnote 212 may also specify positioning information that can be assigned to a set of training images 210. The positioning information for the training images may correspond to the angle of the imaging device that captured the training images, such as the angle of the imaging device relative to the eye or retina.

[0040] In some examples, such annotations 212 for the training image set 210, working distance, or positioning information about the training image set 210, whether correctly or incorrectly positioned, may be assigned manually or through processor-based automation. Furthermore, the training image set 210 may be associated with corresponding visual attributes, such as light distribution parameters, sharpness, brightness, contrast, and other parameters.

[0041] Figures 3 and 4 show a set of training images for training an analysis model, as illustrated in this subject. A set of training images (e.g., set 210) may be stored in a training data repository (e.g., repository 208). In one example, each of the training images in set 210 may correspond to an infrared view of the retina. In this regard, training images may be captured by illuminating the eye with IR light. Although set 210 of training images for training analysis model 204 is shown as IR image frames, such a depiction of set 210 of training images should not be interpreted as limiting. In other examples in this subject, the set of training images may be retinal images, images captured by illuminating the eye with white light, or images captured by illuminating the eye with visible or IR light of any frequency.

[0042] Furthermore, based on the operating conditions, each of the training image set 210 may correspond to either a correctly positioned image or an incorrectly positioned image. Note that correctly positioned images may have the correct operating conditions, thereby fitting the image to the capture, while incorrectly positioned images may have noise or incorrect operating conditions, thereby making the image unsuitable for the capture.

[0043] Figure 3 shows a set of correctly positioned images 300, according to an example. In particular, the set of correctly positioned images 300 can be captured under correct operating conditions. Such correct operating conditions can be achieved, for example, by proper illumination, proper working distance, and proper angle. The set of correctly positioned images 300 is free of any noise, such as white spots, dark boundaries, and unwanted artifacts. Moreover, the set of correctly positioned images 300 accurately captures the eye, specifically the retina, pupil, optic disc, and blood vessels of the eye. For example, the set of correctly positioned images 300 can be captured at the right time, such as when the working distance between the eye and the imaging device is correct, the angle of the imaging device is correct, and the eye or retina is properly illuminated, focused, and captured correctly.

[0044] Figure 4 shows a set of inaccurately positioned images 400, according to an example. In particular, the set of inaccurately positioned images 400 may be captured under incorrect operating conditions. Such incorrect operating conditions can result from, for example, inadequate illumination, incorrect working distance, or inappropriate angle. The set of inaccurately positioned images 400 may have noise such as white spots 402, dark boundaries, and unwanted artifacts 404. Moreover, the set of inaccurately positioned images 400 inaccurately captures the eye, specifically the retina, pupil, optic disc, and blood vessels of the eye. For example, the set of inaccurately positioned images 400 may be captured incorrectly, such as when the working distance between the eye and the imaging device is incorrect, the angle of the imaging device is inappropriate, or the eye or retina is not properly illuminated, i.e., the image frame may be dark, blurry, or out of focus.

[0045] Returning to Figure 2, system 202 may receive a set of training images 210 for training the analysis model 204. In this regard, the training engine 214 of system 202 can train the analysis model 204 based on the set of training images 210. As can be understood, the analysis model 204 incorporates confidence weights 220, which may be further defined during the training process implemented by the training engine 214. The set of confidence weights 220 may refer to learnable parameters of a trainable machine learning model (such as the analysis model 204) defined based on the set of training images 210.

[0046] In one example, the training engine 214 may identify visual attributes for each of the training images 210. As mentioned above, the visual attributes for the training images 210 may include, for example, light distribution parameters, sharpness, brightness, contrast, shape, and structure.

[0047] Subsequently, the training engine 214 may use the analysis model 204 to correlate the visual attributes of the training image set 210 with the corresponding annotations 212 associated with the training image set 210. In one example, the training engine 214 may use the analysis model 204 to correlate the visual attributes of the training images with the annotations associated with the training images, where the annotations indicate whether the training image corresponds to a correctly positioned image or an incorrectly positioned image. The above process may be repeated for each of the training images in the training image set 210. Based on the correlations, the analysis model 204 may understand the visual attributes associated with correctly positioned images and the visual attributes associated with incorrectly positioned images. Based on this understanding, the set of confidence weights 220 of the analysis model 204 may be updated. In one example, the set of confidence weights 220 of the analysis model 204 may be updated based on each of the correlations between the visual attributes of the training image set 210 and the corresponding annotations 212.

[0048] This subject describes training an analysis model 204 using a labeled set of training images 210. However, such supervised training of the analysis model 204 should not be interpreted as limiting. In other examples in this subject, the analysis model 204 may be trained in an unsupervised manner using an unlabeled training dataset. In such cases, the analysis model 204 may, for example, analyze the visual attributes of the set of training images 210 to determine the annotations corresponding to each of the set of training images 210. Accordingly, the set of confidence weights 220 may be updated.

[0049] Once trained, the updated set of confidence weights 220 and the analysis model 204 can be deployed in an evaluation system for automated retinal image capturing. The manner in which the analysis model triggers automated retinal image capturing is described with reference to Figure 5.

[0050] Figure 5 shows an evaluation system 502 implementing a trained analysis model 204, as an example of this subject. The evaluation system 502 (e.g., system 102) includes a processor and memory (not shown) for automated retinal image capturing. The evaluation system 502 may further include the trained analysis model 204. Examples of the evaluation system 502 include, but are not limited to, desktop computers, laptop computers, smartphones, digital cameras, and camcorders.

[0051] Furthermore, the evaluation system 502 may be operably coupled to the imaging device 504. The imaging device 504 may be an electronic device capable of recording, storing, manipulating, and transmitting digital images. The imaging device 504 may include several components, such as an optical aperture, an imaging lens, a microlens array, an imaging element, and a controller. The imaging lens of the imaging device 504 may collect light rays and redirect them towards the imaging element of the imaging device 504. By redirecting the light rays to a single point, an image may be formed on the imaging element. The microlens array of the imaging device 504 may allow the user to adjust the imaging device, such as focusing and zooming. However, although the imaging device 504 is shown to be directly coupled to the evaluation system 502, such depiction should not be interpreted as limiting. In some embodiments of this subject, the evaluation system 502 may be coupled to the imaging device 504 via a network. In one example, the network may be similar to network 206 shown in Figure 2.

[0052] Although the evaluation system 502 is shown as being different from the imaging device 504, it should be noted that, without deviating from the scope of this subject, the evaluation system 502 and the imaging device 504 may be the same device. In this example, the evaluation system 502 may be a smartphone, and the smartphone includes the imaging device 504. In this regard, the imaging device 504 may be the camera of the smartphone.

[0053] The evaluation system 502 may include instructions 506, which, when executed, may implement a trained analysis model 204 for automated retinal image capturing. The retinal image may correspond to a digital picture of the eye in question. The digital picture may show the retina, pupil, optic disc, and blood vessels of the eye in question. In particular, the retina may be a point to which light rays entering the eye can be redirected. As can be understood, an image may be formed on the retina of the eye as light rays enter the eye. Furthermore, the optic disc may be a point on the retina that holds the optic nerve. For example, a retinal image of the retina may be examined to check the health of the eye or to diagnose certain diseases such as macular degeneration, glaucoma, and retinotoxicity.

[0054] The evaluation system 502 may further include a capturing engine 508. The capturing engine 508 may be implemented as a combination of hardware and programming, for example, as programmable instructions for implementing various functionalities. The capturing engine 508 may be implemented in the same manner as the engine 214 described with respect to Figure 2. The capturing engine 508 may perform operations to facilitate retinal image capturing by the evaluation system 502. The system 502 may further include data 510. The data 510 may include a plurality of target frames 512 and other data 514. The other data 514 may be any data generated or used by the evaluation system 502 during its operation.

[0055] During operation, the imaging device 504 may record multiple target frames 512. The multiple target frames 512 may relate to the retina of the subject. The multiple target frames may be recorded in the real-time mode of the imaging device 504. The multiple target frames 512 may correspond to image frames of the field of view of the imaging device 504 in real-time mode. For example, a camera application may be launched on the smartphone to activate the smartphone's camera. The camera may be positioned relative to the subject's eye, for example, in front of or on the retina, in order to record the multiple target frames 512. In such a case, the real-time mode may correspond to the live view of the smartphone's camera, and the multiple target frames 512 may correspond to image frames of the live view. For example, the subject's eye may be illuminated with near-infrared light in order to record the multiple target frames 512 using the imaging device 504.

[0056] The evaluation system 502 may execute a command 506 to receive a plurality of target frames 512 recorded by the imaging device 504. It should be noted that such a plurality of target frames 512 are not retinal images to be captured. The plurality of target frames 512 are interim image frames captured before capturing the retinal image, for example, after launching a camera application on a smartphone. Subsequently, the plurality of target frames 512 may be temporarily stored by the evaluation system 502 for processing.

[0057] The evaluation system 502 may execute instructions 506 to analyze multiple target frames 512 based on the trained analysis model 204. Prior to use in the evaluation system 502, the analysis model 204 may be trained in a remote training system (such as the training system 202 described in relation to Figure 2). As described above, the analysis model 204 may be trained based on a set of training images (such as the set of training images 210). Note that the training images 210 are not retinal images, but rather mediating image frames of the retina of the eye. During training, the analysis model 204 may incorporate a set of confidence weights (such as the set of confidence weights 220) based on the attributes of the set of training images 210. In some examples, the attributes of the set of training images 210 may include, but are not limited to, visual attributes such as light distribution parameters, brightness, contrast, sharpness, shape, structure, and other parameters, as well as annotations. Once the analysis model 204 has been trained for use, it may be deployed on the evaluation system 502.

[0058] In one example, the capturing engine 508 may execute an instruction 506 to analyze each of several target frames 512 based on a trained analysis model 204. Several target frames 512 may be analyzed based on the trained analysis model 204 which incorporates a set of confidence weights 220. In particular, the capturing engine 508 may use the trained analysis model 204 to determine visual attributes of a target frame. In one example, the capturing engine 508 may determine visual attributes at the center and boundary of the target frame, where the center of the target frame may represent the pupil, and the boundary of the target frame may represent the boundary of the retina of the target eye. In this example, the capturing engine 508 may determine visual attributes corresponding to the blood vessels of the target eye.

[0059] In one example, based on the visual attributes of the target frame, the capturing engine 508 may use a trained analysis model 204 to determine the working distance relative to the target frame. The working distance may be the distance between the imaging device 504 and the target retina when the target frame is captured. The capturing engine 508 may use the trained analysis model 204 to further determine positioning information, such as the angle of the imaging device 504 relative to the target retina when the target frame is captured. For example, the analysis model 204 may correlate the determined visual attributes with the working distance corresponding to the target frame.

[0060] The capturing engine 508 can use the analysis model 204 to determine whether the visual attributes of a target frame correspond to a correctly positioned image or an incorrectly positioned image. In one example, the analysis model 204 may determine whether a target frame corresponds to a correctly positioned image or an incorrectly positioned image based on visual attributes, working distance, and positioning information. In this way, multiple target frames 512 can be analyzed. As is understood, multiple target frames 512 may have variations in visual attributes such as light distribution parameters, brightness, contrast, and sharpness. For this purpose, the trained analysis model 204, when implemented on multiple target frames 512, can correlate the visual attributes of the multiple target frames 512 as either correctly positioned or incorrectly positioned images. Based on the analysis, at least one of the multiple target frames may be identified, and the identified target frame may correspond to a correctly positioned image.

[0061] As can be understood, a target frame identified as a correctly positioned image can accurately record the target retina. For example, such a target frame may not be blurry or out of focus, and may not contain any noise such as white spots, unwanted artifacts, and dark boundaries. Furthermore, such a target frame may be recorded from the correct working distance such that the imaging device 504 is appropriately spaced from the target eye and tilted at the correct angle relative to the target eye. As a result, the target frame can accurately represent the target retina.

[0062] The capturing engine 508 may then cause the imaging device 504 to capture a retinal image of the target retina based on an analysis of multiple target frames 512. The capturing engine 508 may trigger the imaging device 504 to automatically capture a retinal image based on the analysis model 204. If the analysis model 204 determines that a target frame corresponds to a correctly positioned image, the capturing engine 508 may trigger the imaging device 504 to capture a retinal image. The retinal image is captured immediately as the target frame corresponding to the correctly positioned image is recorded. The retinal image thus captured has the correct working distance between the imaging device 504 and the retina, and the field of view of the imaging device 504 is properly illuminated, i.e., there are no unwanted artifacts, such as white spots in the center or dark boundary.

[0063] In one example, when triggered, the imaging device 504 may illuminate the target eye with white light before capturing a retinal image. For example, the imaging device 504 may use an electronic flash to illuminate the eye with white light. Once illuminated with white light, the imaging device 504 may capture a retinal image.

[0064] In some cases, the capturing engine 508 may wait to identify at least two consecutive target frames that can correspond to a correctly positioned image. In this regard, the capturing engine 508 may use the analysis model 204 to determine whether a first set of target frames from multiple target frames 512 corresponds to a correctly positioned image. For example, the first set of target frames from multiple target frames 512 includes at least three target frames corresponding to the retina of interest. The capturing engine 508 may then use the analysis model 204 to determine whether the target frames in the first set of target frames are consecutive. For example, the capturing engine 508 may determine whether the target frames in the first set of target frames are consecutive based on the timestamps associated with the first set of target frames. If it determines that the target frames in the first set of target frames should be consecutive, the capturing engine 508 may trigger the imaging device 504 to automatically capture the retinal image.

[0065] This example describes an evaluation system 502 receiving multiple target frames 512 at the beginning of its operation. However, such an implementation of the evaluation system 502 should not be interpreted as limiting. In other implementations of this subject, the reception of multiple target frames 512 may be a continuous process. In this regard, at least one target frame may be received at the beginning of processing or analysis, but other target frames from the multiple target frames 512 may be received at any time. As a result, the multiple target frames 512 may be analyzed as they are acquired.

[0066] Figure 6 shows an exemplary method 600 for training a processor-based model for automated retinal image capture, based on an example of this subject. In some examples, the processor-based model may be an analysis model (e.g., analysis model 204). The order in which the above-described method 600 is presented is not intended to be interpreted as limiting, and some of the method blocks described may be combined in different orders to implement this method or alternative method.

[0067] Furthermore, the method 600 described above may be implemented with appropriate hardware, computer-readable instructions, or a combination thereof. The steps of such a method may be carried out by a system under instructions from machine-executable instructions stored on a non-temporary computer-readable medium, or by dedicated hardware circuits, microcontrollers, or logic circuits. For example, the method 600 may be carried out by the training system 202. In this specification, some examples cover non-temporary computer-readable mediums that are computer-readable, such as digital data storage media, and are also intended to encode computer-executable instructions, the instructions which carry out some or all of the steps of the method described above.

[0068] In block 602, a set of training images is received. The set of training images (e.g., set 210) may correspond to intermediate image frames of the retina. In one example, the training system 202 may receive set 210 of training images from a repository source, such as the training data repository 208. In another example, a user may upload set 210 of training images to the training system 202.

[0069] In one example, each of the training image set 210 may have a corresponding attribute. Such attributes may include, for example, visual attributes and other parameters. Additionally, the training image set 210 may be associated with a corresponding annotation (such as an annotation from annotation 212) on which the analysis model 204 can be trained. For example, an annotation about a training image may indicate that the training image is either correctly positioned or incorrectly positioned.

[0070] In block 604, the processor-based model is trained on a set of training images. System 202 may include an analysis model 204. In one example, a training engine (such as training engine 214) may execute instructions to train the analysis model 204 on a set of training images 210. Attributes of the set of training images 210, such as visual attributes and other parameters, as well as annotations 212 associated with the set of training images 210, may be considered for training the analysis model 204. For example, each visual attribute of the set of training images 210 may be correlated with a corresponding annotation from annotations 212. Based on the correlation, the analysis model 204 may be trained to identify visual attributes for correctly positioned images and visual attributes for incorrectly positioned images.

[0071] In block 606, the set of confidence weights for the processor-based model may be updated based on training. In one example, the analysis model 204 incorporates a set of confidence weights 220. For example, the set of confidence weights 220 may represent the correlation between the visual attributes of the training image set 210 and the corresponding annotations 212. During the training of the analysis model 204, the set of confidence weights 220 incorporated into the analysis model 204 may be refined and updated by the training engine 212.

[0072] In block 608, a trained processor-based model incorporating an updated set of confidence weights may be sent to an evaluation system for automated retinal image capturing. For this purpose, a trained analysis model 204 incorporating an updated set of confidence weights 220 may be sent to an evaluation system (such as evaluation system 502). In evaluation system 502, the analysis model 204 may be run by an engine, such as a capturing engine 508, to analyze multiple target frames 512. The capturing engine 508 can use the trained analysis model 204 to determine the correctly positioned target frame from the multiple target frames 512. As a result, the capturing engine 508 can trigger the imaging device 504 for automated capture of the retinal image of the target retina.

[0073] Figure 7 shows a computing environment 700 that implements a non-temporary computer-readable medium for training an analysis model for automated retinal image capturing. In one example, the computing environment 700 includes a processor 702 that is communicably coupled to the non-temporary computer-readable medium 704 via a communication link 706. In one example, the processor 702 may have one or more processing resources for fetching and executing computer-readable instructions from the non-temporary computer-readable medium 704. The processor 702 and the non-temporary computer-readable medium 704 may be implemented, for example, in a system 202 (described in relation to Figure 2).

[0074] The non-temporary computer-readable medium 704 may be, for example, an internal memory device or an external memory device. In the exemplary implementation, the communication link 706 may be a network communication link. The processor 702 and the non-temporary computer-readable medium 704 may be communicated via a network to a training data repository 708 (similar to the training data repository 208). The processor 702 and the non-temporary computer-readable medium 704 may similarly be communicated via a network to an evaluation system (such as the evaluation system 502).

[0075] In an exemplary implementation, the non-temporary computer-readable medium 704 includes a set of computer-readable instructions 710 that can be accessed by the processor 702 through a communication link 706. Referring to Figure 7, in one example, the non-temporary computer-readable medium 704 includes instructions 710 that cause the processor 702 to receive a set of training images, such as a set of training images 210. In one example, the set of training images 210 may be intermediary image frames corresponding to an infrared view of the retina captured by an imaging device. Instructions 710 may cause the processor 702 to train an analysis model 204 based on the set of training images 210. In one example, the set of training images 210 may have corresponding visual attributes, such as light distribution parameters, brightness, contrast, sharpness, and other parameters, and may be associated with corresponding annotations. For training the analysis model 204, the visual attributes for each of the set of training images 210 may be correlated with and examined in relation to the corresponding annotations. In one example, visual attributes corresponding to correctly positioned images and visual attributes corresponding to incorrectly positioned images can be determined from a set of training images 210.

[0076] During the training of the analysis model 204, instruction 710 may further cause the processor 702 to update the set of confidence weights for the analysis model 204. The analysis model 204 may incorporate a set of confidence weights 220. For example, a training engine (such as training engine 214) may trigger periodic updates of the set of confidence weights 220 while training the analysis model. Note that the updated values ​​of the set of confidence weights 220 may be refined or improved compared to the old values ​​of the set of confidence weights 220.

[0077] After training the analysis model 204, an instruction 710 may be executed to cause the processor 702 to send the trained analysis model 204, incorporating an updated set of confidence weights 220, to the system. The analysis model 204 may also incorporate visual attributes of correctly positioned and incorrectly positioned images. The trained analysis model 204, sent to the system (e.g., evaluation system 502), may be executed by a capturing engine (e.g., capturing engine 508). The capturing engine 508 may analyze multiple target frames corresponding to the target retina based on the trained analysis model 204. Based on the analysis, the capturing engine 508 may determine whether a target frame from the multiple target frames corresponds to a correctly positioned image. Once it has determined that a target frame corresponds to a correctly positioned image, the capturing engine 508 may cause an imaging device (e.g., imaging device 504) to capture a retinal image of the target retina.

[0078] While examples for this disclosure are described using terminology specific to structural features and / or methods, it should be understood that the attached claims are not necessarily limited to the specific features or methods described. Rather, specific features and methods are disclosed and described as examples of this disclosure. [Explanation of Symbols]

[0079] 102 System 104 Processors 106 Machine-readable storage medium 202 Training System, System 204 Analysis Models 206 Network 208 Training data repositories, repositories 214 Training Engine, Engine 502 Evaluation system, system 504 Imaging devices 508 Capturing Engine 700 Computing Environments 702 Processor 704 Non-temporary computer-readable media 706 Communication Link 708 Training Data Repository

Claims

1. A system for automated retinal image capture, wherein the system is Processor and The system comprises a machine-readable storage medium containing instructions, the instructions being: Receiving multiple target frames from an imaging device, wherein the multiple target frames relate to the target retina, and the multiple target frames are recorded in real-time mode by the imaging device. Analyzing the multiple target frames based on an analysis model, The analysis model is trained on a set of training images, the analysis model incorporates a set of confidence weights, and the set of confidence weights analyzes based on the visual attributes and annotations for each of the training images. Based on the analysis of the plurality of target frames, the imaging device is made to capture a retinal image of the target retina. A system that can be executed by the processor to perform the following actions.

2. The system according to claim 1, wherein the visual attributes relating to the training image correspond to visual attributes across the retinal boundary in the training image, visual attributes at the center of the retina, visual attributes of the pupil of the eye in the training image, and visual attributes of the blood vessels inside the eye.

3. The system according to claim 1, wherein the visual attributes of the training image include one of the following in the training image: light distribution parameters, sharpness, contrast, and brightness.

4. The system according to claim 1, wherein annotations for the training image indicate the working distance between the imaging device and the retina in the training image, and positioning information of the imaging device.

5. The system according to claim 1, wherein annotations for training images indicate that the training images correspond to either correctly positioned images suitable for capture or incorrectly positioned images unsuitable for capture.

6. The analysis of the aforementioned multiple target frames is as follows: Determining the visual attributes for each of the aforementioned multiple target frames, Determining whether at least one of the plurality of target frames corresponds to a correctly positioned image suitable for capture, based on the corresponding visual attributes, The system according to claim 1, further comprising causing the imaging device to capture the retinal image of the retina based on the aforementioned determination.

7. The analysis of the aforementioned multiple target frames is as follows: To determine whether a first set of target frames from the plurality of target frames corresponds to a correctly positioned image, To determine whether the target frames within the first set of target frames are consecutive, The system according to claim 6, further comprising, based on the above determination, causing the imaging device to capture the retinal image of the retina.

8. The system according to claim 7, wherein the first set of target frames from the plurality of target frames includes at least three target frames corresponding to the retina of the target.

9. The system according to claim 1, wherein the retina is illuminated with near-infrared light in order to capture the plurality of target frames using the imaging device.

10. The system according to claim 6, wherein when it is determined that at least one of the plurality of target frames corresponds to a correctly positioned image that fits the capture, the retina is illuminated with white light to cause the capturing of the retinal image.

11. A method for training a processor-based model for automated retinal image capture, A step of receiving a set of training images, each of which corresponds to an infrared view of the retina, A step of training the processor-based model based on the set of training images, The steps include updating the set of confidence weights of the processor-based model based on the aforementioned training, The steps include sending the trained processor-based model incorporating the updated set of confidence weights to a system for automated retinal image capture, and Methods that include...

12. The step of training the aforementioned processor-based model is: The steps include determining the visual attributes for each of the aforementioned training images, The method according to claim 11, further comprising the step of correlating the visual attributes for each of the set of training images with corresponding annotations.

13. The method according to claim 12, wherein annotations for training images indicate that the training images correspond to correctly positioned images suitable for capture, or to incorrectly positioned images unsuitable for capture.

14. The method according to claim 11, wherein the processor-based model is trained in a training system.

15. The method according to claim 11, wherein the processor-based model, once trained, is deployed on an evaluation system operably coupled to an imaging device.

16. The evaluation system is, The receiving of multiple target frames from the imaging device, wherein the multiple target frames relate to the target retina, and the multiple target frames are recorded in real-time mode by the imaging device. Based on the aforementioned trained processor-based model, the plurality of target frames are analyzed, Based on the analysis of the plurality of target frames, the imaging device is made to capture a retinal image of the target retina. The method according to claim 15, wherein the method is performed.

17. A non-temporary computer-readable medium comprising instructions, wherein the instructions are: Receiving a set of training images, each of which corresponds to the infrared view of the retina, Training an analysis model based on the aforementioned set of training images, Based on the aforementioned training, the set of confidence weights for the analysis model is updated, Send the trained analysis model incorporating the updated set of confidence weights to the system. A non-temporary, computer-readable medium that can be executed by processing resources to perform the following actions.

18. In the aforementioned system, receiving the trained analysis model, In the aforementioned system, the system receives a plurality of target frames relating to the target retina from an imaging device, wherein the plurality of target frames are recorded in real-time mode by the imaging device. In the system, based on the trained analysis model, the plurality of target frames are analyzed to determine whether at least one of the plurality of target frames corresponds to a correctly positioned image that fits the capture. The system causes the imaging device to capture a retinal image of the target retina based on the determination. A non-temporary computer-readable medium according to claim 17, comprising instructions that can be executed by processing resources to perform the following:

Citation Information

Patent Citations

  • Assessment of fundus images

    CN111345775A

  • Assessment of fundus images

    EP3669754A1

  • Smartphone-based handheld optical device and method for capturing non-mydriatic retinal images

    EP3695775A1

  • Ophthalmologic apparatus, control method of ophthalmologic apparatus, and program

    JP2021097989A

  • Automated fundus imaging system

    US7458685B2