Generating a visible light image from a scanning laser ophthalmoscope image

A deep learning-based model converts SLO images into visible light equivalents, addressing interpretability issues and enhancing diagnostic capabilities by providing familiar representations for clinicians, thus improving retinal examination efficiency and accessibility.

WO2025191613A1PCT designated stage Publication Date: 2025-09-18REMIDIO INNOVATIVE SOLUTIONS PRIVATE LIMITED
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/IN2025/050363
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-13
Filing Date
2025-03-13
Publication Date
2025-09-18

AI Technical Summary

Technical Problem

Scanning Laser Ophthalmoscopy (SLO) images are complex and difficult for clinicians to interpret due to their infrared and pseudo-color nature, limiting their usability in routine clinical practice, while traditional visible light images are preferred for ease of diagnosis.

Method used

A deep learning-based image generation model, trained on paired SLO and visible light images, converts SLO images into visible light equivalents by leveraging generative adversarial networks (GANs) to extract and translate retinal features, providing a familiar representation for clinicians.

Benefits of technology

The model effectively transforms SLO images into visible light images, enhancing diagnostic capabilities by offering a wider field of view and improving accessibility of retinal examinations, facilitating telemedicine applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IN2025050363_18092025_PF_FP_ABST
    Figure IN2025050363_18092025_PF_FP_ABST
Patent Text Reader

Abstract

Approaches for generating a visible light image from a scanning laser ophthalmoscope (SLO) image are described. In an example, an input SLO image of a retina is obtained and processed using an image generation model. This model is trained on paired datasets of SLO images and corresponding visible light images, along with detailed retinal feature data. The model extracts retinal features from the input SLO image, including information about various retinal layers, structures, and characteristics. Based on these extracted features, the model generates a visible light image that closely resembles what would be captured by a traditional visible light retinal camera. The image generation model utilizes a generative adversarial network (GAN) architecture, comprising a generator and discriminator, to produce highly realistic synthetic images.
Need to check novelty before this filing date? Find Prior Art

Description

GENERATING A VISIBLE LIGHT IMAGE FROM A SCANNING LASER OPHTHALMOSCOPE IMAGEBACKGROUND

[0001] Eyes are the human body's most highly developed sensory organs, playing a crucial role in our perception and interaction with the world. They are essential for vision, light detection, and contribute significantly to overall health. Regular examination and imaging of the eyes are vital for early detection, diagnosis, and monitoring of various ocular and systemic conditions, potentially preventing vision loss and other complications. Retinal imaging is a key tool in ophthalmology, providing detailed views of the eye's internal structures. These images are critical for diagnosing and tracking the progression of various eye diseases, as well as identifying signs of systemic conditions that may manifest in the retina. Different imaging techniques have been developed to capture detailed views of the eye's structures, each with its own advantages.BRIEF DESCRIPTION OF FIGURES

[0002] Systems and / or methods, is accordance with examples of the present subject matter are now described and with reference to the accompanying figures, in which:

[0003] FIGS. 1 A-1 B illustrates a training system for training a image generation model, as per an example;

[0004] FIG. 2 illustrates an image generation system for generating a visible light image based on a scanning laser ophthalmoscope (SLO) image, as per another example;

[0005] FIG. 3 illustrates a set of training images used in the image generation process, as per an example.

[0006] FIG. 4 illustrates the effectiveness of the image generation model, showing (A) an aligned pseudo-color SLO image and (B) the corresponding visible light image generated by an image generation model, as per an example;

[0007] FIG. 5 illustrates a method for training an image generation model, as per an example;

[0008] FIG. 6 illustrates a method for generating a visible light image based on a scanning laser ophthalmoscope (SLO) image, using a trained image generation model, as per an example; and

[0009] FIG. 7 illustrates a system environment implementing a non- transitory computer readable medium for generating a visible light image based on a scanning laser ophthalmoscope (SLO) image using an image generation model, as per an example.DETAILED DESCRIPTION

[0010] Eyes are important sensory organs for various reasons, including vision, detecting light, and maintaining overall health. Imaging the eyes helps in diagnosing and monitoring a range of ocular and systemic conditions. One effective method for imaging the eyes is capturing retinal images using visible light cameras. However, during this imaging process, the pupil may act as a barrier. When the pupil constricts, it limits the amount of light entering the eye, which in turn restricts the field of view captured by the camera. This may be particularly challenging for patients with naturally small pupils or those sensitive to light, as it may obscure important details needed for accurate diagnosis.

[0011] In another example, an alternative method for imaging the retina of the eyes is using Scanning Laser Ophthalmoscopy (SLO). SLO works by scanning the retina with a low-power laser beam, which allows for high- resolution imaging. Instead of using a camera sensor to capture light reflected from the retina, laser beams scan the retina one retinal point at a time. This technique allows for a wider field of view for the same pupil size and better imaging of patients with lens opacity, such as cataracts. SLO overcomes the drawback of the pupil acting as a barrier because it uses a confocal pinhole to eliminate out-of-focus light, thereby enhancing image clarity even when the pupil is constricted. Additionally, SLO may captureimages through smaller pupils and can be used in low-light conditions, making it suitable for patients with light sensitivity or naturally small pupils.

[0012] However, SLO images have their own drawbacks. The resulting images do not consist of traditional visible light images. Instead, they include an infrared image, a red-free image, and a pseudo-color image. The pseudo-color image is a composite created by merging visible light wavelengths. This complexity makes it hard for clinicians to interpret the images, as they often require advanced interpretation skills. The high- resolution images may sometimes highlight artifacts or irrelevant details, potentially complicating the diagnostic process. Additionally, the specialized nature of SLO images can limit their accessibility and usability in routine clinical practice.

[0013] Approaches for generating a visible light image from a scanning laser ophthalmoscope (SLO) image, are described. The SLO image may be an image of the eye which is captured using one or more specific wavelengths of light, such as infrared (IR), red-free (RF), or a combination thereof to produce a pseudo-color image. Such SLO image may be either stored in a data repository or may be captured by a SLO. In one example, SLO image, corresponding to a retina of a subject eye, which is to be used for generating a visible light image, is obtained. The SLO image may contain detailed information about the retinal structure, including blood vessels, optic disc, macula, and other features that are typically visualized in ophthalmological examinations. This SLO image serves as the input for a deep learning-based system that aims to generate a visible light equivalent, providing clinicians with a familiar representation of the retina that closely resembles images captured by traditional visible light retinal cameras.

[0014] Once obtained, the SLO image is processed using an image generation model. This model is a sophisticated deep learning algorithm that has been trained on a diverse dataset comprising pairs of training SLO images and corresponding training visible light images captured by a visible light camera. The training process enables the model to learn the complexrelationships between the features present in SLO images and their visible light counterparts. The training dataset may include a wide range of retinal conditions and variations to ensure the model's robustness and generalizability. During the training phase, the model may utilize techniques such as image registration to ensure precise alignment between the SLO and visible light image pairs. The image generation model may be trained based on advanced neural network architectures, such as generative adversarial networks (GANs) or U-Net variants, to achieve results in image- to-image translation tasks.

[0015] The image generation model performs two primary functions. First, it extracts retinal features from the input SLO image, identifying and characterizing structural and characteristic details of the retina of the subject's eye. These features may include the intricate network of blood vessels, the shape and condition of the optic disc, the appearance of the macula, and other clinically relevant structures. Second, based on these extracted retinal features, the model generates an output visible light image of the retina. This visible light image aims to closely mimic the appearance and characteristics of an image that would have been captured by a traditional visible light retinal camera, providing clinicians with a familiar and interpretable representation of the retinal structures.

[0016] The present invention offers significant technical advancements in the field of retinal imaging and diagnosis. By leveraging deep learning techniques, it enables the conversion of Scanning Laser Ophthalmoscopy (SLO) images into visible light equivalents, bridging the gap between advanced imaging technologies and traditional clinical interpretation. This approach overcomes limitations associated with pupil size and allows for a wider field of view in retinal examinations. The system's ability to extract and translate complex retinal features from SLO data to familiar visible light representations enhances diagnostic capabilities without requiring additional imaging procedures. Furthermore, the use of a trained image generation model allows for real-time processing and visualization,potentially improving the efficiency and accessibility of retinal examinations. This technology may also facilitate telemedicine applications by providing standardized, easily interpretable retinal images regardless of the capture method, thus expanding the reach of specialized ophthalmic care to underserved areas.

[0017] The manner in which sub-models implemented within the image generation model are trained and used for generating the visible light image based on the SLO image is explained in detail with respect to FIGS. 1 -4. While aspects of described systems may be implemented in any number of different electronic devices, environments, and / or implementation, the examples are described in the context of the following example device(s). In another example, the aspects of the present subject matter may also be implemented by a standalone device having executable instructions. It may be noted that drawings of the present subject matter shown here are for illustrative purposes and are not to be construed as limiting the scope of the subject matter claimed.

[0018] FIG. 1 A illustrates a training system 102 comprising a processor or memory (not shown), for training an image generation model or submodels included within the image generation model. In an example, the training system 102 (referred to as system 102) may be communicatively coupled to a repository 104 through a network 106. The repository 104 may further include training data 108. The training data 108 may include data that may be used for training the image generation model. In an example, the training data 108 includes training SLO images and corresponding training visible light images for training the image generation model.

[0019] In another example, the training data 108 may include detailed retinal feature data corresponding to each training visible light image. This retinal feature data may comprise information about various retinal structures and characteristics, such as the nerve fiber layer, ganglion cell layer, inner plexiform layer, inner nuclear layer, outer plexiform layer, outer nuclear layer, and retinal pigment epithelium. It may also include data onreflectance characteristics, choroidal and scleral layer properties, optical aberrations, eye movements, and contrast and resolution details. These features are carefully annotated and aligned with the corresponding training visible light images to provide comprehensive training information for the image generation model. Although depicted as being obtained from a single repository, such as repository 104, the training data 108 may also be obtained from multiple other sources without deviating from the scope of the present subject matter. In such cases, each of such multiple repositories may be interconnected through a network, such as network 106.

[0020] The network 106 may be a private network or a public network and may be implemented as a wired network, a wireless network, or a combination of a wired and wireless network. The network 106 may also include a collection of individual networks, interconnected with each other and functioning as a single large network, such as the Internet. Examples of such individual networks include, but are not limited to, Global System for Mobile Communication (GSM) network, Universal Mobile Telecommunications System (UMTS) network, Personal Communications Service (PCS) network, Time Division Multiple Access (TDMA) network, Code Division Multiple Access (CDMA) network, Next Generation Network (NGN), Public Switched Telephone Network (PSTN), Long Term Evolution (LTE), and Integrated Services Digital Network (ISDN).

[0021] The system 102 may further include instructions 1 10 and a training engine 1 12. In an example, the instructions 110 are fetched from a memory and executed by a processor included within the system 102. The training engine 1 12 may be implemented as a combination of hardware and programming, for example, programmable instructions to implement a variety of functionalities. In examples described herein, such combinations of hardware and programming may be implemented in several different ways. For example, the programming for the training engine 1 12 may be executable instructions, such as instructions 1 10. Such instructions may be stored on a non-transitory machine-readable storage medium which may becoupled either directly with the system 102 or indirectly (for example, through networked means). In an example, the training engine 112 may include a processing resource, for example, either a single processor or a combination of multiple processors, to execute such instructions. In the present examples, the non-transitory machine-readable storage medium may store instructions, such as instructions 110, that when executed by the processing resource, implement training engine 1 12. In other examples, the training engine 1 12 may be implemented as electronic circuitry.

[0022] The instructions 110, when executed by the processing resource, cause the training engine 112 to train the image generation model 1 14 based on the training data 108. The system 102 may further include a training SLO image(s) 1 16, a training retinal feature data 1 18, and a training visible light image(s) 120. In an example, the system 102 may obtain training data 108 corresponding to a single pair of training images from the repository 104, and the data pertaining to that is stored as training SLO image(s) 1 16, training retinal feature data 1 18, and training visible light image(s) 120.

[0023] As described previously, the image generation model 1 14 may include sub-models including a generator model and a discriminator model. An example of such machine learning models includes deep learning models. Although the present examples have been described in relation to deep learning models, the aforementioned approaches may also be implemented using other machine learning models. It may also be noted that any explanation provided in conjunction with deep learning models is applicable to other machine learning models, without limitations and without deviating from the scope of the present subject matter. Such examples have not been described for the sake of brevity. The manner in which the training of the generator model and the discriminator model may be performed is further described in conjunction with FIG. 1 B.

[0024] FIG. 1 B illustrates a block diagram of the image generation model 1 14, which is to be trained for generating visible light images basedon SLO images. The image generation model 114 employs a generative adversarial network (GAN) architecture, including two primary sub-models 122, i.e., a generator model 124 and a discriminator model 126. The generator model 124 is the key component of the image generation model 1 14 and is trained to generate the visible light image based on input SLO image. In an example, the SLO image may be infrared, red-free, or pseudocolor images captured using specific wavelengths of light. The discriminator 126 is the second component of the image generation model 1 14 and is trained to generate an output indicating differences between the image generated by the generator model 124, i.e., the visible light image, and a real visible light image, which acts as a reference image. The output of the discriminator model 126 is then provided as feedback to the generator model 124 to refine its training process.

[0025] In operation, the training engine 1 12 obtains training data 108 including training SLO image(s) 1 16 and corresponding training visible light image(s) 120. In an example, the training SLO image(s) 1 16 are captured by a SLO and the training visible light image(s) 120 are captured by a visible light camera. These image pairs represent the same retinal areas captured using different modalities. The training SLO images 1 16 may include infrared images captured at wavelengths of 700-800 nm, red-free images captured at wavelengths of 500-600 nm, pseudo-color images obtained by merging infrared and red-free images, or combination thereof. Each training SLO image is paired with a corresponding training visible light image of the same retina, ensuring a diverse dataset that covers various retinal conditions and appearances.

[0026] In one example, the training data 108, including training SLO image(s) 1 16 and training visible light image(s) 120, may be augmented to increase the size and diversity of the dataset. This augmentation process employs various techniques to create additional training samples. These techniques may include random brightness changes within a range of 0.1 , random contrast adjustments between 0.9 and 1.1 , and the addition ofrandom noise within a 0.01 range. Additionally, the system 102 may apply geometric transformations such as horizontal flipping, vertical flipping, and 90-degree rotation, each with a 0.5 probability of being applied to a given image. These data augmentation techniques help to improve the robustness and generalization capabilities of the image generation model 1 14 by exposing it to a wider variety of image variations during the training process.

[0027] In addition, the training engine 112 may obtain training retinal feature data 1 18 corresponding to a plurality of retinal features. This comprehensive data includes detailed information about retinal layers, such as the nerve fiber layer, ganglion cell layer, inner plexiform layer, inner nuclear layer, outer plexiform layer, outer nuclear layer, and retinal pigment epithelium. The feature data also encompasses reflectance characteristics of different retinal tissues, properties of choroidal and scleral layers, optical aberrations characteristics that may affect image quality, eye movements during image capture, and contrast and resolution details. This rich feature set provides the model with in-depth knowledge of retinal structures and their appearance across different imaging modalities.

[0028] Once the training data 108 is obtained, the training engine 1 12 aligns the training SLO image(s) 1 16 with the corresponding training visible light image(s) 1 18. In an example, the alignment of the training images included within the training data 108 is performed in two stages. Initially, during the dataset curation process, the training engine 1 12 aligns the training SLO image(s) 1 16 with the corresponding training visible light image(s) 120 using advanced image registration techniques. These techniques may include both rigid and non-rigid transformations to account for potential distortions between the two imaging modalities. Image pairs for which this initial alignment process fails are excluded from the training dataset.

[0029] Subsequently, during the training process itself, a fine-tuning of the alignment is performed using a Spatial Transformer Network (STN). The STN is a differentiable module that may be inserted into the neural networkarchitecture or within the system 202 (not shown in FIG. 1 ), allowing the model to learn and apply spatial transformations to input images. It may perform operations such as scaling, cropping, rotations, and non-rigid deformations, dynamically adjusting the alignment of features between the SLO and visible light images. This adaptive alignment capability of the STN helps compensate for any residual misalignments and variations in image geometry that may persist after the initial registration step. This two-step alignment process ensures that the same retinal structures appear at corresponding pixel locations in both the training SLO and training visible light images, facilitating accurate learning of the relationship between SLO and visible light representations of retinal structures. The precise alignment is crucial for the image generation model 1 14 to learn the correct mapping between the different imaging modalities.

[0030] Thereafter, the training engine 1 12 trains the image generation model 114 based on the training data 108 which is aligned. As described above, the image generation model 114 is implemented as a Generative Adversarial Network (GAN) architecture, including the generator model 124 and the discriminator model 126. The training engine 1 12 first trains the generator model 124 to produce visible light images corresponding to the input SLO images. The generator model 124, which may be based on a U- Net architecture or similar deep convolutional network, learns to extract relevant features from the input SLO images (depicted by arrow 128 in FIG. 1 B) and translate them into the visible light domain images (depicted by arrow 130). It progressively learns to map the unique characteristics of SLO images, such as the high contrast of blood vessels in infrared images, to their appearance in visible light.

[0031] Concurrently or sequentially, the training engine 112 trains the discriminator model 126 to distinguish between real visible light images and the output visible light images generated by the generator model 124. The discriminator model 126, typically implemented as a convolutional neuralnetwork, learns to identify subtle differences between real and synthetic images, pushing the generator model 124 to produce more realistic results.

[0032] Once both sub-models 122 are initially trained, the training engine 1 12 operates the generator model 124 to generate visible light images 130 from input SLO images 128 and provides these as input to the discriminator model 126. The discriminator model 126 then attempts to distinguish between the visible light images 130 and real visible light images 132, generating an output 134 indicating the perceived differences. This output is then provided as feedback 136 to the generator model 124, allowing it to refine its image generation process. The generator model 124 uses this feedback 136 to adjust its parameters, aiming to produce images that may better fool the discriminator model 126. This adversarial training process continues iteratively, with the generator model 124 improving its ability to create realistic visible light images from SLO inputs, while the discriminator model 126 becomes more adept at detecting subtle inconsistencies.

[0033] The training process continues until both the generator model 124 and the discriminator model 126 reach a state of equilibrium, where neither is able to significantly improve its performance relative to the other. This balance is achieved when both models have optimized their respective tasks - the generator model 124 in producing realistic visible light images from SLO inputs, and the discriminator model 126 in distinguishing between real and generated images. At this point, further training would not yield substantial improvements in the quality of the generated images or the discriminator's ability to differentiate them. The generator model 124 is then considered to be sufficiently trained to effectively transform SLO images into visible light images that closely resemble those captured by visible light retinal cameras, preserving important retinal features and structures while presenting them in a familiar visible light format.

[0034] The trained generator model 124 may then be deployed independently to generate visible light equivalents corresponding to newSLO images, potentially enhancing the diagnostic capabilities of SLO imaging by providing outputs in a format that clinicians are more accustomed to interpreting.

[0035] The manner in which the image generation model 1 14 may be used for generating visible light images based on input SLO images is further described in conjunction with FIG. 2.

[0036] FIG. 2 illustrates an environment 200 with an image generation system 202 for generating an output visible light image based on an input SLO image. In an example, the image generation system 202 (referred to as system 202) includes a mobile phone, tablet, mac-book, or any other portable computing device having a camera device coupled to it. The system 202 obtains the input SLO image from a data repository 204 which is connected over a network (not shown in Figure) with the system 202. The data repository 204 further includes a plurality of input SLO image(s) 206 which may be obtained for generating visible light images. In another example, the system 202 may obtain SLO images from a SLO device which may be captured in real-time. Further, the system 202 may also be implemented within the SLO itself to obtain the SLO image and generate visible light image based on SLO image.

[0037] The system 202 may include a processor 208, interface(s) 210, and memory(s) 212. The processor 208 may be implemented as microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, state machines, logic circuitries, and / or other devices that manipulate signals based on operational instructions. Among other capabilities, the processor 208 may be configured to obtain various types of data, such as flight data, operational parameters, and weather data. The processor 208 may then use an image generation model, such as image generation model 1 14, to generate visible light image based on the input SLO image.

[0038] The interface(s) 210 may allow the connection or coupling of the system 202 with one or more sensors or devices onboard the system 202,depending on the implementation of the system 202 through a wired network, a wireless network, or a combination of a wired and wireless network. The interface(s) 210 may also enable intercommunication between different logical as well as hardware components of the system 202.

[0039] The memory(s) 212 may be a computer-readable medium, examples of which include volatile memory (e.g., RAM), and / or non-volatile memory (e.g., Erasable Programmable read-only memory, i.e., EPROM, flash memory, etc.). The memory(s) 212 may be an external memory, or internal memory, such as a flash drive, a compact disk drive, an external hard disk driver, or the like. The memory(s) 212 may further include data which either may be utilized or generated during the operation of the system 202.

[0040] Similar to the system 102, the system 202 may further include instruction(s) 214 and engine(s) 216. In an example, the instruction(s) 214 are fetched from the memory(s) 212 and executed by the processor 208 included within the system 202. The engine(s) 216 may include image generation engine 218 and other engine(s) 220. The other engine(s) 220 may further implement functionalities that supplement functions performed by the system 202 or any of the engine(s) 216. The image generation engine 218 (referred to as engine 218) may be implemented as a combination of hardware and programming, for example, programmable instructions to implement a variety of functionalities. In examples described herein, such combinations of hardware and programming may be implemented in several different ways. For example, the programming for the engine 218 may be executable instructions, such as instruction(s) 214. Such instruction(s) 214 may be stored on a non-transitory machine-readable storage medium which may be coupled either directly with the system 202 or indirectly (for example, through networked means). In an example, the engine 218 may include a processing resource, for example, either a single processor or a combination of multiple processors, to execute such instructions. In the present examples, the non-transitory machine-readable storage mediummay store instructions, such as instruction(s) 214, that when executed by the processing resource, implement the engine 218. In other examples, the engine 218 may be implemented as electronic circuitry.

[0041] The system 202 may further include an image generation model, such as image generation model 114, and a data 222. The data 222 may include corresponding data that is utilized or generated by the system 202, while performing a variety of functions. In an example, the data 222 further includes input SLO image(s) 224, retinal feature(s) 226, output visible light image(s) 228, and other data 230. Further, the other data 230, amongst other things, may serve as a repository for storing data that is processed, or received, or generated as a result of the execution of the instructions by the processor 208.

[0042] In operation, an input SLO image of a retina of a subject's eye, which needs to be processed to generate a visible light image for the same is obtained. In an example, the engine 218 within the image generation system 202 obtains the input SLO image(s) 206 from the data repository 204. These input SLO image(s) 206 may be stored as the input SLO image(s) 224 within the data 222. The input SLO image(s) 224 may be one of an infrared image, a red-free image, a pseudo-color image, or a combination thereof. In some cases, the pseudo-color image may be obtained by merging the infrared image with the red-free image. The wavelength of light used to capture the infrared images may be in the range of 700-800 nm, while the wavelength for red-free images may be in the range of 500-600 nm.

[0043] Once obtained, the engine 218 uses the image generation model 1 14 to extract retinal features from the input SLO image(s) 224. The image generation model 1 14, implemented within the system 202, processes the input SLO image(s) 224 to identify and characterize various structural and characteristic details of the retina. These features may include information about retinal layers such as the nerve fiber layer, ganglion cell layer, inner plexiform layer, inner nuclear layer, outer plexiform layer, outer nuclearlayer, and retinal pigment epithelium. The model also extracts data on reflectance, choroidal and scleral layers characteristics, optical aberrations characteristics, eye movements, and contrast and resolution details. This comprehensive set of extracted features provides a detailed representation of the retinal structure as captured in the input SLO image(s) 224.

[0044] As described in conjunction with FIG. 1A-1 B, the image generation model 1 14 is trained based on a diverse dataset of paired SLO and visible light images, along with corresponding retinal feature data. This training process enables the image generation model 1 14 to learn the complex relationships between SLO image characteristics and their visible light counterparts. The image generation model 1 14 comprises a pair of sub-models: the generator model 124 and the discriminator model 126, trained using a generative adversarial network (GAN) architecture. During training, the generator model 124 learns to produce visible light images corresponding to input SLO images, while the discriminator model 126 learns to distinguish between real and generated visible light images. This adversarial training process results in a highly capable model that may accurately transform SLO images into realistic visible light equivalents.

[0045] Continuing with the present example, based on the extracted retinal feature data, which is temporarily stored in the retinal feature(s) 226, the engine 218 generates the output visible light image(s) 228 of the retina. For example, the generator component of the image generation model 1 14 takes the extracted features as input and synthesizes a visible light image that closely resembles what would be captured by a visible light retinal camera. This process involves translating the unique characteristics of SLO images, such as the high contrast of blood vessels in infrared images, into their appearance in visible light. The generated output visible light image(s) 228 preserves the important structural and pathological features of the retina while presenting them in a format that is more familiar to clinicians accustomed to interpreting visible light retinal images.

[0046] In another example, the system 202 may be communicatively coupled to a central computing server through a network (not shown in FIG. 2). The network may be a private network or a public network and may be implemented as a wired network, a wireless network, or a combination of a wired and wireless network, and may be similar to the network 106 (as depicted in FIG. 1 A). All the above disclosed steps which may be performed by the engine 218 of the system 202, may be implemented or performed by the central computing server on behalf of the system 202 to reduce computing load on edge of the network.

[0047] FIG. 3 illustrates a set of training images used in the image generation process. The figure is arranged in a 2x2 grid, with each quadrant labeled from (A) to (D). Image (A) shows an infrared training SLO image, typically captured at wavelengths between 700-800 nm. This image highlights deep retinal structures and the choroidal vasculature. Image (B) presents a red-free training SLO image, usually captured at wavelengths of 500-600 nm, which emphasizes retinal nerve fiber layer and surface blood vessels. Image (C) displays a pseudo-color SLO image, created by merging the infrared and red-free images, providing a composite view that combines the benefits of both imaging modalities. Finally, image (D) shows a real visible light image of the same retinal area, serving as the ground truth for training the image generation model 1 14. This arrangement allows for direct comparison between different SLO imaging techniques and the target visible light image.

[0048] FIG. 4 illustrates images demonstrating the effectiveness of the image generation model. For example, image (A) shows a real visible light image that has been carefully aligned with the pseudo-color image, e.g., aligned with image (C) of FIG. 3. This alignment process ensures that retinal structures in the SLO image precisely match their positions in the visible light domain. Image (B) presents the output generated by the image generation model. This visible light image is created from the input SLO image, showcasing the model's ability to transform SLO data into a formatclosely resembling traditional visible light retinal photography. The side-by- side comparison allows for evaluation of the model's performance in preserving important retinal features while translating them into the visible light spectrum.

[0049] FIG. 5 illustrates an example method 500 for training an image generation model, in accordance with examples of the present subject matter. The order in which the above-mentioned method is described is not intended to be construed as a limitation, and some of the described method blocks may be combined in a different order to implement the method, or alternative method.

[0050] Furthermore, the above-mentioned method may be implemented in a suitable hardware, computer-readable instructions, or combination thereof. The steps of such method may be performed by either a system under the instruction of machine executable instructions stored on a non- transitory computer readable medium or by dedicated hardware circuits, microcontrollers, or logic circuits. For example, the method may be performed by a training system, such as system 102. In an implementation, the method may be performed under an “as a service” delivery model, where the system 102, operated by a provider, receives programmable code. Herein, some examples are also intended to cover non-transitory computer readable medium, for example, digital data storage media, which are computer readable and encode computer-executable instructions, where said instructions perform some or all the steps of the above- mentioned method.

[0051] In an example, the method 500 may be implemented by the system 102 for training an image generation model based on a training data such as training data 108. At block 502, a training data including a training SLO image and a training visible light image is obtained. For example, the training engine 1 12 obtains training data 108 including training SLO image(s) 1 16 and corresponding training visible light image(s) 120. In an example, the training SLO image(s) 1 16 are captured by a SLO and thetraining visible light image(s) 120 are captured by a visible light camera. These image pairs represent the same retinal areas captured using different modalities. The training SLO images 116 may include infrared images captured at wavelengths of 700-800 nm, red-free images captured at wavelengths of 500-600 nm, pseudo-color images obtained by merging infrared and red-free images, or combination thereof. Each training SLO image is paired with a corresponding training visible light image of the same retina, ensuring a diverse dataset that covers various retinal conditions and appearances.

[0052] At block 504, the training SLO image is aligned pixel by pixel with the training visible light image. For example, the training engine 1 12 aligns the training SLO image(s) 1 16 with the corresponding training visible light image(s) 1 18. In an example, the alignment of the training images included within the training data 108 is performed in two stages. Initially, during the dataset curation process, the training engine 1 12 aligns the training SLO image(s) 116 with the corresponding training visible light image(s) 120 using advanced image registration techniques. Subsequently, during the training process itself, a fine-tuning of the alignment is performed using the STN. The precise alignment is crucial for the model to learn the correct mapping between the different imaging modalities.

[0053] At block 506, a generator model is trained to produce an output visible light image corresponding to the input SLO image. For example, the training engine 112 trains the image generation model 1 14 based on the training data 108 which is aligned. As described above, the image generation model 1 14 is implemented as a Generative Adversarial Network (GAN) architecture, including the generator model 124 and the discriminator model 126. The training engine 112 first trains the generator model 124 to produce visible light images corresponding to the input SLO images. The generator model 124, which may be based on a U-Net architecture or similar deep convolutional network, learns to extract relevant features from the input SLO images (depicted by arrow 128 in FIG. 1 B) and translate theminto the visible light domain images (depicted by arrow 130). It progressively learns to map the unique characteristics of SLO images, such as the high contrast of blood vessels in infrared images, to their appearance in visible light.

[0054] At block 508, a discriminator model is trained to distinguish between a real visible light image and the output visible light image generated by the generator model. For example, the training engine 1 12 trains the discriminator model 126 to distinguish between real visible light images and the output visible light images generated by the generator model 124. The discriminator model 126, typically implemented as a convolutional neural network, learns to identify subtle differences between real and synthetic images, pushing the generator model 124 to produce more realistic results.

[0055] At block 510, the discriminator model is used to generate an output indicating differences between real visible light image and the output visible light image. For example, once both sub-models 122 are initially trained, the training engine 1 12 operates the generator model 124 to generate visible light images 130 from input SLO images 128 and provides these as input to the discriminator model 126. The discriminator model 126 then attempts to distinguish between the visible light images 130 and real visible light images 132, generating an output 134 indicating the perceived differences.

[0056] At block 512, the output of the discriminator model is provided as feedback to the generator model to refine the training of the generator model. For example, the output generated by the discriminator model 126 is then provided as feedback 136 to the generator model 124, allowing it to refine its image generation process. The generator model 124 uses this feedback 136 to adjust its parameters, aiming to produce images that may better fool the discriminator model 126. This adversarial training process continues iteratively, with the generator model 124 improving its ability tocreate realistic visible light images from SLO inputs, while the discriminator model 126 becomes more adept at detecting subtle inconsistencies.

[0057] The training process continues until both the generator model 124 and the discriminator model 126 reach a state of equilibrium, where neither is able to significantly improve its performance relative to the other. This balance is achieved when both models have optimized their respective tasks - the generator model 124 in producing realistic visible light images from SLO inputs, and the discriminator model 126 in distinguishing between real and generated images. At this point, further training would not yield substantial improvements in the quality of the generated images or the discriminator's ability to differentiate them. The generator model 124 is then considered to be sufficiently trained to effectively transform SLO images into visible light images that closely resemble those captured by visible light retinal cameras, preserving important retinal features and structures while presenting them in a familiar visible light format.

[0058] The trained generator model 124 may then be deployed independently to generate visible light equivalents corresponding to new SLO images, potentially enhancing the diagnostic capabilities of SLO imaging by providing outputs in a format that clinicians are more accustomed to interpreting.

[0059] FIG. 6 illustrates example method 600 for generating a visible light image for an input SLO image using a trained image generation model. Similar to FIG. 3, the order in which the above-mentioned method is described is not intended to be construed as a limitation, and some of the described method blocks may be combined in a different order to implement the method, or alternative method. Based on the present approaches as described in the context of the example method 600, the input SLO image(s) 224 are processed based on the trained image generation model 1 14.

[0060] Further, the above-mentioned method 600 may be implemented in a suitable hardware, computer-readable instructions, or combination thereof. The steps of such method may be performed by either a systemunder the instruction of machine executable instructions stored on a non- transitory computer readable medium or by dedicated hardware circuits, microcontrollers, or logic circuits. For example, the method may be performed by an image generation system, such as system 202. In an implementation, the method may be performed under an “as a service” delivery model, where the system 202, operated by a provider, receives programmable code. Herein, some examples are also intended to cover non-transitory computer readable medium, for example, digital data storage media, which are computer readable and encode computer-executable instructions, where said instructions perform some or all the steps of the above-mentioned method.

[0061] In an example, the method 600 may be implemented by the system 202 for generating a visible light image for an input SLO image using the trained image generation model 114. At block 602, an input SLO image of a retina of a subject’s eye is obtained. For example, the engine 218 within the image generation system 202 obtains the input SLO image(s) 206 from the data repository 204. These input SLO image(s) 206 may be stored as the input SLO image(s) 224 within the data 222. The input SLO image(s) 224 may be one of an infrared image, a red-free image, a pseudo-color image, or a combination thereof. In some cases, the pseudo-color image may be obtained by merging the infrared image with the red-free image. The wavelength of light used to capture the infrared images may be in the range of 700-800 nm, while the wavelength for red-free images may be in the range of 500-600 nm.

[0062] At block 604, an image generation model is used to extract retinal features from the input SLO image. For example, the engine 218 uses the image generation model 1 14 to extract retinal features from the input SLO image(s) 224. The image generation model 1 14, implemented within the system 202, processes the input SLO image(s) 224 to identify and characterize various structural and characteristic details of the retina. These features may include information about retinal layers such as the nerve fiberlayer, ganglion cell layer, inner plexiform layer, inner nuclear layer, outer plexiform layer, outer nuclear layer, and retinal pigment epithelium. The model also extracts data on reflectance, choroidal and scleral layers characteristics, optical aberrations characteristics, eye movements, and contrast and resolution details. This comprehensive set of extracted features provides a detailed representation of the retinal structure as captured in the input SLO image(s) 224.

[0063] At block 606, the image generation model is used to generate an output visible light image of the retina based on the extracted retinal features. For example, based on the extracted retinal feature data, which is temporarily stored in the retinal feature(s) 226, the engine 218 generates the output visible light image(s) 228 of the retina. For example, the generator component of the image generation model 1 14 takes the extracted features as input and synthesizes a visible light image that closely resembles what would be captured by a visible light retinal camera. This process involves translating the unique characteristics of SLO images, such as the high contrast of blood vessels in infrared images, into their appearance in visible light. The generated output visible light image(s) 228 preserves the important structural and pathological features of the retina while presenting them in a format that is more familiar to clinicians accustomed to interpreting visible light retinal images.

[0064] FIG. 7 illustrates a computing environment 700 implementing a non-transitory computer-readable medium for generating visible light images from scanning laser ophthalmoscope (SLO) images. The computing environment 700 includes processor(s) 702 communicatively coupled to a computer-readable medium 704 through a communication link 706. The processor(s) 702 may have one or more processing resources for fetching and executing computer-readable instructions from the computer-readable medium 704.

[0065] The computer-readable medium 704 may be, for example, an internal memory device or an external memory device. In an exampleimplementation, the communication link 706 may be a network communication link. The processor(s) 702 and the computer-readable medium 704 may also be communicatively coupled to a client device 708 over the network.

[0066] In an example implementation, the computer-readable medium 704 includes a set of computer-readable instructions 710 (referred to as instructions 710) which may be accessed by the processor(s) 702 through the communication link 706. The instructions 710 cause the processor(s) 702 to obtain an input SLO image, such as input SLO image(s) 224 of a retina of a subject's eye, and use an image generation model, such as image generation model 1 14, to process the input SLO image. The image generation model 1 14 is trained based on training data 108 comprising a training SLO image and a training visible light image captured by a visible light camera.

[0067] The input SLO image(s) 224 may be one of an infrared image, a red-free image, a pseudo-color image, or combination thereof. The pseudocolor image may be obtained by merging the infrared image with the red- free image. Typically, the wavelength of light used to capture the infrared images is in the range of 700-800 nm, while for red-free images, it's in the range of 500-600 nm.

[0068] In an example, the image generation model is trained based on training retinal feature data comprising data corresponding to a plurality of retinal features. These features include details about retinal layers such as the nerve fiber layer, ganglion cell layer, inner plexiform layer, inner nuclear layer, outer plexiform layer, outer nuclear layer, and retinal pigment epithelium. It also includes data on reflectance characteristics, choroidal and scleral layers characteristics, optical aberrations characteristics, eye movements, and contrast and resolution details.

[0069] Continuing further, the instructions 710 cause the processor(s) 702 to use the trained image generation model 114 to extract retinal feature(s) 226 from the input SLO image(s) 224, indicating structural andcharacteristic details of the retina of the subject's eye. The processor(s) 702 then generates an output visible light image, such as output visible light image(s) 228 of the retina based on the extracted retinal feature(s) 226.

[0070] As described above as well, the image generation model 1 14 includes a pair of sub-models, namely, the generator model 124 and the discriminator model 126, trained using the GAN architecture. The generator model 124 is trained to produce visible light images corresponding to the input SLO images, while the discriminator model 126 is trained to distinguish between real visible light images and the output visible light images generated by the generator model 124.

[0071] During the training process, the discriminator model 126 generates an output indicating the differences between the real visible light image and the output visible light image. This output is used as feedback to refine the generator's performance, allowing it to produce increasingly realistic synthetic visible light images.

[0072] The instructions 710 enable the system to process various types of SLO images and generate corresponding visible light images. These images preserve the important structural and pathological features of the retina while presenting them in a format that is more familiar to clinicians accustomed to interpreting visible light retinal images. This capability may potentially enhance the diagnostic value of SLO imaging by providing outputs in a format that clinicians are more accustomed to interpreting.

[0073] Although examples for the present disclosure have been described in language specific to structural features and / or methods, it is to be understood that the appended claims are not necessarily limited to the specific features or methods described. Rather, the specific features and methods are disclosed and explained as examples of the present disclosure.

Claims

l / We Claim:1 . A system comprising: a processor; and an image generation engine coupled to the processor, wherein the image generation engine is to: obtain an input scanning laser ophthalmoscope (SLO) image of a retina of a subject’s eye; use an image generation model, wherein the image generation model is trained based on training data comprising a training SLO image and a training visible light image captured by a visible light camera, and wherein the image generation model, is to: extract retinal features from the input SLO image indicating structural and characteristic details of the retina of the subject’s eye; and generate an output visible light image of the retina based on the extracted retinal features.

2. The system as claimed in claim 1 , wherein the image generation model is trained based on training retinal feature data comprising data corresponding to a plurality of retinal features, wherein the retinal features comprise details about retinal layers, comprising the nerve fiber layer, ganglion cell layer, inner plexiform layer, inner nuclear layer, outer plexiform layer, outer nuclear layer, and retinal pigment epithelium, reflectance characteristics, choroidal and scleral layers characteristics, optical aberrations characteristics, eye movements, and contrast and resolution details.

3. The system as claim in claim 1 , wherein the input SLO image is one of an infrared image, a red-free image, a pseudo-color image, orcombination thereof, wherein the pseudo color image is obtained by merging the infrared image with the red free image.

4. The system as claimed in claim 3, wherein the wavelength of light used to capture the infrared images is 700-800 nm and the wavelength of light used to capture the red-free images is 500-600 nm.

5. The system as claimed in claim 1 , wherein the image generation model comprises a pair of sub-models comprising a generator and a discriminator, trained using a generative adversarial network (GAN) architecture.

6. The system as claimed in claim 5, wherein the generator is trained to produce visible light image corresponding to the input SLO image and the discriminator is trained to distinguish between a real visible light image and the output visible light image generated by the generator, wherein the discriminator generates an output indicating the differences between the real visible light image and the output visible light image.

7. A method comprising: obtaining training data comprising a training SLO image and a training visible light image captured by a visible light camera; training an image generation model based on training data, wherein the image generation model, when trained, to: extract features from an input SLO image of a retina, wherein the input SLO image corresponds to a subject eye captured by a SLO; and generate an output visible light image of the retina based on the extracted features.

8. The method as claimed in claim 7, wherein the method comprises:obtaining training retinal feature data comprising data corresponding to a plurality of retinal features for training the image generation model, wherein the retinal features comprise details about retinal layers, comprising the nerve fiber layer, ganglion cell layer, inner plexiform layer, inner nuclear layer, outer plexiform layer, outer nuclear layer, and retinal pigment epithelium, reflectance characteristics, choroidal and scleral layers characteristics, optical aberrations characteristics, eye movements, and contrast and resolution details.

9. The method as claimed in claim 7, wherein the method further comprises: aligning the training SLO image pixel by pixel with the training visible light image, wherein the alignment of training images is performed to achieve same retina structure at same pixel location in both the training images.

10. The method as claimed in claim 7, wherein the training SLO image is one of infrared image, red-free image, pseudo-color image, or combination thereof.1 1 . The method as claimed in claim 7, wherein the image generation model comprises a pair of sub-models comprising a generator model and a discriminator model, trained using Generative Adversarial Network (GAN) architecture.

12. The method as claimed in claim 11 , wherein the training the image generation model comprises: training the generator to produce visible light image corresponding to the input SLO image.

13. The method as claimed in claim 11 , wherein the training the image generation model comprises: training a discriminator to distinguish between a real visible light image and the output visible light image generated by the generator.

14. The method as claimed in claim 11 , wherein the training the image generation model comprises: using the discriminator to generate an output indicating differences between real visible light image and the output visible light image; and providing the output of the discriminator as feedback to the generator to refine the training of the generator.

15. A non-transitory computer-readable medium comprising instructions, the instructions being executable by a processing resource to: obtain an input scanning laser ophthalmoscope (SLO) image of a retina of a subject’s eye; use an image generation model, wherein the image generation model is trained based on training data comprising a training SLO image and a training visible light image captured by a visible light camera, and wherein the image generation model is to: extract retinal features from the input SLO image indicating structural and characteristic details of the retina of the subject’s eye; and generate an output visible light image of the retina based on the extracted retinal features.

16. The non-transitory computer-readable medium of claim 15, wherein the image generation model is trained based on training retinal feature data comprising data corresponding to a plurality of retinal features, wherein the retinal features comprise details about retinal layers, comprising the nerve fiber layer, ganglion cell layer, inner plexiform layer, inner nuclear layer,outer plexiform layer, outer nuclear layer, and retinal pigment epithelium, reflectance characteristics, choroidal and scleral layers characteristics, optical aberrations characteristics, eye movements, and contrast and resolution details.

17. The non-transitory computer-readable medium of claim 15, wherein the input SLO image is one of an infrared image, a red-free image, a pseudo color image, or combination thereof, wherein the pseudo color image is obtained by merging the infrared image with the red free image.

18. The non-transitory computer-readable medium of claim 17, wherein the wavelength of light used to capture the infrared images is 700-800 nm and the wavelength of light used to capture the red-free images is 500-600 nm.

19. The non-transitory computer-readable medium of claim 15, wherein the image generation model comprises a pair of sub-models comprising a generator and a discriminator, trained using a generative adversarial network (GAN) architecture.

20. The non-transitory computer-readable medium of claim 19, wherein the generator is trained to produce visible light image corresponding to the input SLO image and the discriminator is trained to distinguish between a real visible light image and the output visible light image generated by the generator, wherein the discriminator generates an output indicating the differences between the real visible light image and the output visible light image.

Citation Information

Patent Citations

  • Scanning laser ophthalmoscope laser guidance for laser vitreolysis

    CA3234401A1

  • Pattern analysis of retinal maps for the diagnosis of optic nerve diseases by optical coherence tomography

    US20080309881A1

  • Optical texture analysis of the inner retina

    US20190110681A1