Comprehensive ai-powered cross-platform multimodal ophthalmic data management, analysis, and prediction
An AI-driven system for generating predicted retinal images addresses the limitation of conventional techniques by enabling early detection and proactive treatment of retinal conditions through advanced image analysis and registration.
Patent Information
- Application Number
- PCT/US2025/013981
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-12-18
- Filing Date
- 2025-01-31
- Publication Date
- 2025-08-07
AI Technical Summary
Conventional techniques for analyzing retinal changes are limited in identifying adverse conditions until they become noticeable, limiting treatment options and efficacy as conditions progress.
A system utilizing AI-driven, multimodal ophthalmic image management and analysis software that generates predicted images of patient retinas using machine-learning architectures, capable of detecting disease signatures and biomarkers from OCT images, and performing image registration across various modalities.
Enables early identification of retinal conditions, improving treatment timing and reducing the need for invasive interventions by predicting adverse conditions before they become symptomatic.
Smart Images

Figure US2025013981_07082025_PF_FP_ABST
Abstract
Description
COMPREHENSIVE AI-POWERED CROSS-PLATFORM MULTIMODAL OPHTHALMIC DATA MANAGEMENT, ANALYSIS, AND PREDICTIONSTATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT
[0001] This invention was made with government support under EY008098 awarded by the National Institutes of Health. The government has certain rights in the invention.CROSS-REFERENCE TO RELATED TECHNOLOGY
[0002] This application claims priority to U.S. Provisional Application No. 63 / 549,278, filed February 2, 2024, and International Patent Application No. PCT / US2024 / 060693 filed December 18, 2024, each of which is hereby incorporated by reference in its entirety.TECHNICAL FIELD
[0003] This application generally relates to techniques for generating predicted images of patient retinas and, in some embodiments, to techniques for generating images for patient retinas based on images of patient retinas captured at multiple points in time.BACKGROUND
[0004] Clinicians (e.g., optometrists and ophthalmologists) may use ophthalmoscopes or retinal cameras to look at the posterior portion of the patient’s eyeball. This can be done over a period of visits by the patient. In cases where the clinician identifies certain changes to the eyeball over time indicative of possible issues (e.g., retinal detachment, macular degeneration, etc.), further testing may be performed to confirm whether one or more conditions are present.
[0005] But these conventional techniques are limited in that potentially adverse conditions must first be present so as to be identifiable by the clinician. For example, in the case of retinal detachment, the clinician must first either receive indications from the patient that they are experiencing symptoms such as floaters, fixed shadows, etc. Alternatively, clinicians must visually identify retinal tears or subtle hemorrhage beneath the retina. Addressing these adverse conditions as they occur can limit the available treatment options and the efficacy of such options as the conditions progress.SUMMARY
[0006] Embodiments described herein include systems and methods for improving upon shortcomings and the art and may provide any number of additional or alternative benefits as well. Embodiments generally include systems and methods for implementing artificial intelligence (Al) driven, multimodal, ophthalmic image management, registration, and automated analysis software tool for enhancing diagnostic precision and clinical decision-making. Embodiments include computing hardware and software for obtaining and analyzing image data for aspects of human eyeballs (e g., layers of the eyeball, retina characteristics) for medical study and treatment. A computing system includes software for executing software programming for one or more machine-learning architectures having various machine-learning models, which may include various artificial intelligence (Al) models that receive input data generated at optical coherence tomography (OCT) devices. The computing system may receive an OCT image from an OCT device or from a database and input the OCT image to the Al model. The Al model is trained to generate and output a segmented image from the OCT image. The Al model is trained using any number of example image pairs as training data, though such example image pairs are not needed for actual usage by users (e g., clinicians, patients, researchers, or other types of users) following training. The training data may include training labels indicating expected attributes or aspects of the training images or image pairs. In some implementations, the computing system and machinelearning architecture may include certain types of machine-learning models (e.g., Pix2Pix model) and / or other types of preprocessing operations on the OCT image data to improve low-quality inputted OCT image. As an example, the computing system may execute a Pix2Pix model and image preprocessing operations on a low-quality spectral-domain OCT (SD-OCT) image. Generally, embodiments include a computing system that obtains data associated with a first image (e g., OCT images described herein) captured by an OCT device while imaging a retina of a patient, generates a segmented image based on the first image, determines an image pair based on the first image and the segmented image, and executes a model using the image pair to cause the model to generate an output. The output of the model includes a predicted image of the retina of the patient.
[0007] Embodiments may further include a computing system executing software functions of an ophthalmic image processing tool to detect the presence of disease signatures and age-related macular degeneration. For instance, the computing system may detect and segmentinstances of drusen in a diverse patient population receiving ophthalmic care. The computing system executes software implementing a machine-learning architecture trained to identify OCT images that contain drusenoid lesions and mark those lesions. During training, the machinelearning models of the machine-learning architecture may be trained to be robust against image artifacts and noise. The computing system may calculate biomarkers based on segmented regions of the OCT that can be used for the prediction and evaluation of disease progression and treatment response in age-related macular degeneration.
[0008] Embodiments may further include a computing system executing software functions of an ophthalmic image processing tool to identify and analyze choroid layer (CL) information from OCT images, which may be beneficial for screening and management of choroidal and retinal diseases. The computing system generates or otherwise facilitates biomarker quantification for various downstream operations of disease screening tools. Generating these biomarker quantification outputs may include, for example, using the OCT scans to generate CL thickness maps, CL vascularity index (CVI) maps, CL inner / outer surface contour maps, CL area / volume measurements, CL vessel volume measurements, CL vessel diameter measurements, and CL vessel diameter heatmaps. In some cases, generating these biomarker quantification outputs may include generating area / volume measurements of various retinal lesions including subretinal fluid (SRF), pigment epithelium detachment (PED), and intra-retinal fluid using OCT scans. In some cases, generating these biomarker quantification outputs may include generating measurements of retinal lesions including hard exudates using color fundus (CF) photographs. Additionally or alternatively, generating these biomarker quantification outputs may include executing software programming of a multimodal manual measurement tool for measuring thickness, area, and volume of lesions, as well as various layers of the posterior segment of the eye, among other types of outputs. The executable software may perform multimodal image registration including, for example, registration across various 2D retinal image modalities; (ii) registration of longitudinal 2D retinal images obtained from the same modality; (iii) volumetric registration of longitudinal volumes; (iv) registration of 2D image modalities such as CF, fundus autofluorescence (FAF) with 3D OCT volumes; and (v) registration of 2D and 3D retinal images with Humphrey visual field data.
[0009] Embodiments may include systems including a computing device comprising at least one processor, configured to: receive data associated with a first image and a second image, the first image captured by an optical coherence tomography (OCT) device at a first point in time and the second image captured by the OCT device at a second point in time while imaging a retina of a patient; generate a first segmented image based on the first image and a second segmented image based on the second image, the first segmented image and the second segmented image associated with at least one boundary line denoting a boundary between a first layer of the retina and a second layer of the retina; determine an image pair including the first image and the second image based on the first segmented image and the second segmented image; and execute a model using the image pair to cause the model to generate an output representing a predicted image of the retina of the patient.
[0010] The computer may determine a correspondence between the first image and the second image based on comparing the first segmented image and the second segmented image.
[0011] The first point in time may be chronologically earlier than the second point in time. The first point in time may be at least a year before the second point in time.
[0012] When providing the image pair to the model, the computer may be programmed to execute a generative adversarial network (GAN) using the image pair.
[0013] The model may be a generative adversarial network (GAN). When executing the GAN using the image pair, the computer may be programmed to: execute the model using the image pair to train the model to generate the output including data associated with the predicted image.
[0014] The computer may be further configured to match a histogram of the first image and the second image to a predetermined histogram.
[0015] Embodiments may include a computer-implemented method for retina imaging, including: receiving, by at least one processor, data associated with a first image and a second image, the first image captured by an optical coherence tomography (OCT) device at a first point in time and the second image captured by the OCT device at a second point in time while imaging a retina of a patient; generating, by the at least one processor, a first segmented image based onthe first image and a second segmented image based on the second image, the first segmented image and the second segmented image associated with at least one boundary line denoting a boundary between a first layer of the retina and a second layer of the retina; determining, by the at least one processor, an image pair including the first image and the second image based on the first segmented image and the second segmented image; and executing, by the at least one processor, a model using the image pair to cause the model to generate an output representing a predicted image of the retina of the patient.
[0016] The computer may determine a correspondence between the first image and the second image based on comparing the first segmented image and the second segmented image.
[0017] The first point in time may be chronologically earlier than the second point in time. The first point in time may at least a year before the second point in time.
[0018] When providing the image pair to the model, the computer may execute a generative adversarial network (GAN) using the image pair.
[0019] The model may be a generative adversarial network (GAN). When executing the GAN using the image pair, the computer may execute the GAN using the image pair to train the GAN to generate the output including data associated with the predicted image.
[0020] The computer may match a histogram of the first image and the second image to a predetermined histogram.
[0021] Embodiments may include a non-transitory computer-readable medium storing instructions for retina imaging. When executed by at least one processor, the instructions cause the at least one processor to: receive data associated with a first image and a second image, the first image captured by an optical coherence tomography (OCT) device at a first point in time and the second image captured by the OCT device at a second point in time while imaging a retina of a patient; generate a first segmented image based on the first image and a second segmented image based on the second image, the first segmented image and the second segmented image associated with at least one boundary line denoting a boundary between a first layer of the retina and a second layer of the retina; determine an image pair including the first image and the second image based on the first segmented image and the second segmented image; and execute a model using theimage pair to cause the model to generate an output representing a predicted image of the retina of the patient.
[0022] The instructions may further cause the at least one processor to determine a correspondence between the first image and the second image based on comparing the first segmented image and the second segmented image. The first point in time may be chronologically earlier than the second point in time. The first point in time may be at least a year before the second point in time.
[0023] When the instructions cause the at least one processor to provide the image pair to the model, the instructions may further cause the at least one processor to execute a generative adversarial network (GAN) using the image pair.
[0024] The model may be a generative adversarial network (GAN). When the instructions that cause the at least one processor to execute the GAN using the image pair, the instructions may further cause the at least one processor to execute the GAN using the image pair to train the GAN to generate the output including data associated with the predicted image.
[0025] Embodiments may include a computer-implemented method for disease signature detection in optical coherence tomography (OCT) imaging, the method including: obtaining, by a computer, volumetric data representing three-dimensional imagery of an eyeball based upon a plurality of cross-sectional images received from an OCT imaging device; identifying, by the computer, one or more layer boundaries in the plurality of cross-sectional images of the volumetric data, the one or more layer boundaries corresponding to one or more retinal layers of the eyeball; generating, by the computer, an en-face image representing two-dimensional imagery at a retinal layer corresponding to a layer boundary in the volumetric data; determining, by the computer, one or more features of the en-face image corresponding to one or more disease indictors based on execution of a segmentation model and the en-face image, the segmentation model trained to receive the en-face image as input, identify one or more disease indicators in the one or more features, and generate an output including a segmented en-face image; and generating, by the computer, an updated en-face image having the segmented en-face image having the one or more disease indicators identified in the input en-face image.
[0026] The computer identifies a layer boundary corresponding to at least one of a retinal pigment epithelium (RPE) layer or a vascular layer. The computer may identify an instance of a drusen structure in the one or more disease indicators of the one or more features at the RPE layer. The computer may identify an instance of geographic atrophy (GA) structure in the one or more disease indicators of the one or more features at the vascular layer.
[0027] The plurality of cross-sectional images of the volumetric data may represent two- dimensional imagery at a first axis of the eyeball from the OCT imaging device. The en-face image may represent two-dimensional imagery at a second axis of the eyeball at the retinal layer of the volumetric data.
[0028] When generating the en-face image, for each cross-sectional image of the plurality of cross-section images, the computer may determine a reflectivity value of one or more pixels of the cross-sectional image. The computer may generate an average reflectivity value using each reflectivity value of the plurality cross-section images according to the second axis. A disease indicator of the one or more features may include a reflectivity value satisfying a segmentation threshold.
[0029] Embodiments may include a system for disease signature detection in optical coherence tomography (OCT) imaging. The system may include: a computer including at least one processor, configured to: obtain volumetric data representing three-dimensional imagery of an eyeball based upon a plurality of cross-sectional images received from an OCT imaging device; identify one or more layer boundaries in the plurality of cross-sectional images of the volumetric data, the one or more layer boundaries corresponding to one or more retinal layers of the eyeball; generate an en-face image representing two-dimensional imagery at a retinal layer corresponding to a layer boundary in the volumetric data; determine one or more features of the en-face image corresponding to one or more disease indictors based on execution of a segmentation model and the en-face image, the segmentation model trained to receive the en-face image as input, identify one or more disease indicators in the one or more features, and generate an output including a segmented en-face image; and generate an updated en-face image having the segmented en-face image having the one or more disease indicators identified in the input en-face image.
[0030] The computer identifies a layer boundary corresponding to at least one of a retinal pigment epithelium (RPE) layer or a vascular layer. The computer may identify an instance of a drusen structure in the one or more disease indicators of the one or more features at the RPE layer. The computer may identify an instance of geographic atrophy (GA) structure in the one or more disease indicators of the one or more features at the vascular layer.
[0031] The plurality of cross-sectional images of the volumetric data may represent two- dimensional imagery at a first axis of the eyeball from the OCT imaging device. The en-face image may represent two-dimensional imagery at a second axis of the eyeball at the retinal layer of the volumetric data.
[0032] When generating the en-face image, for each cross-sectional image of the plurality of cross-section images, the computer may be configured to determine a reflectivity value of one or more pixels of the cross-sectional image. The computer may be configured to generate an average reflectivity value using each reflectivity value of the plurality cross-section images according to the second axis. A disease indicator of the one or more features may include a reflectivity value satisfying a segmentation threshold.
[0033] Embodiments may include a non-transitory computer-readable medium storing instructions for disease signature detection in optical coherence tomography (OCT) imaging. When executed by at least one processor, the instructions may cause the at least one processor to: obtain volumetric data representing three-dimensional imagery of an eyeball based upon a plurality of cross-sectional images received from an OCT imaging device; identify one or more layer boundaries in the plurality of cross-sectional images of the volumetric data, the one or more layer boundaries corresponding to one or more retinal layers of the eyeball; generate an en-face image representing two-dimensional imagery at a retinal layer corresponding to a layer boundary in the volumetric data; determine one or more features of the en-face image corresponding to one or more disease indictors based on execution of a segmentation model and the en-face image, the segmentation model trained to receive the en-face image as input, identify one or more disease indicators in the one or more features, and generate an output including a segmented en-face image; and generate an updated en-face image having the segmented en-face image having the one or more disease indicators identified in the input en-face image.
[0034] The computer identifies a layer boundary corresponding to at least one of a retinal pigment epithelium (RPE) layer or a vascular layer. The computer may identify an instance of a drusen structure in the one or more disease indicators of the one or more features at the RPE layer. The computer may identify an instance of geographic atrophy (GA) structure in the one or more disease indicators of the one or more features at the vascular layer.
[0035] The plurality of cross-sectional images of the volumetric data may represent two- dimensional imagery at a first axis of the eyeball from the OCT imaging device. The en-face image may represent two-dimensional imagery at a second axis of the eyeball at the retinal layer of the volumetric data.
[0036] When generating the en-face image, for each cross-sectional image of the plurality of cross-section images, the computer may determine a reflectivity value of one or more pixels of the cross-sectional image. The computer may generate an average reflectivity value using each reflectivity value of the plurality cross-section images according to the second axis.
[0037] Embodiments may include computer-implemented method for choroid analysis of tomography (OCT) imaging, the method including: obtaining, by a computer, volumetric data representing three-dimensional imagery of an eyeball based upon a plurality of cross-sectional images received from an OCT imaging device; identifying, by the computer, one or more layer boundaries corresponding to a choroid layer of the eyeball based on the plurality of cross-sectional images of the volumetric data; identifying, by the computer, a plurality of vessel boundaries corresponding to the plurality of vessels of the choroid layer of the eyeball based on the plurality of cross-sectional images of the volumetric data; generating, by the computer, a three-dimensional representation of each vessel identified in the choroid layer of the eyeball; and generating, by the computer, a heatmap based on the three-dimensional representation of each vessel, the heatmap representing variation in diameter of the plurality of vessels the eyeball along centerlines defined by the plurality of vessels.
[0038] For each vessel in the choroid layer, the instructions may further cause the at least one processor to determine a diameter based upon point cloud data of a point cloud representing the vessel of the eyeball. The instructions may further cause the at least one processor to generate the heatmap based upon each diameter determined for each vessel of the plurality of vessels.
[0039] When generating the three-dimensional representation of each vessel identified in the choroid layer, for each vessel of the choroid layer of the eyeball, instructions may further cause the at least one processor to identify a centerline for the vessel as represented by at least one cross- sectional image. The instructions may further cause the at least one processor to determine a plurality of cross-sectional lines extending along the centerline for the vessel. The instructions may further cause the at least one processor to generate the three-dimensional representation of the vessel as part of a point cloud representing the eyeball based on the centerline and the plurality of cross-sectional lines.
[0040] For each vessel of the choroid layer, the instructions may further cause the at least one processor to execute a tensor voting operation for determining the centerline and a centerline point for the vessel.
[0041] The instructions may further cause the at least one processor to determine a plurality of vessel metrics for the plurality of vessels of the eyeball using the three-dimensional representation of each vessel in the choroid layer, including a diameter for each particular vessel of the plurality of vessels. The instructions may further cause the at least one processor to generate an output file associated with the heatmap, the output file including the plurality of vessel metrics of the plurality of vessels.
[0042] When generating the three-dimensional representation of each vessel identified in the choroid layer of the eyeball, the instructions may further cause the at least one processor to generate a vector image representation of each vessel identified in the choroid layer of the eyeball.
[0043] Embodiments may include a system for choroid analysis of tomography (OCT) imaging, the system including: a computer including at least one processor, configured to: obtain volumetric data representing three-dimensional imagery of an eyeball based upon a plurality of cross-sectional images received from an OCT imaging device; identify one or more layer boundaries corresponding to a choroid layer of the eyeball based on the plurality of cross-sectional images of the volumetric data; identify a plurality of vessel boundaries corresponding to the plurality of vessels of the choroid layer of the eyeball based on the plurality of cross-sectional images of the volumetric data; generate a three-dimensional representation of each vessel identified in the choroid layer of the eyeball; and generate a heatmap based on the three-dimensional representation of each vessel, the heatmap representing variation in diameter of the plurality of vessels the eyeball along centerlines defined by the plurality of vessels.
[0044] For each vessel in the choroid layer, the computer may determine a diameter based upon point cloud data of a point cloud representing the vessel of the eyeball. The computer may generate the heatmap based upon each diameter determined for each vessel of the plurality of vessels.
[0045] When generating the three-dimensional representation of each vessel identified in the choroid layer, for each vessel of the choroid layer of the eyeball, the computer may identify a centerline for the vessel as represented by at least one cross-sectional image. The computer may determine a plurality of cross-sectional lines extending along the centerline for the vessel. The computer may generate the three-dimensional representation of the vessel as part of a point cloud representing the eyeball based on the centerline and the plurality of cross-sectional lines.
[0046] For each vessel of the choroid layer, the computer may execute a tensor voting operation for determining the centerline and a centerline point for the vessel.
[0047] The computer may determine a plurality of vessel metrics for the plurality of vessels of the eyeball using the three-dimensional representation of each vessel in the choroid layer, including a diameter for each particular vessel of the plurality of vessels. The computer may generate an output fde associated with the heatmap, the output file including the plurality of vessel metrics of the plurality of vessels.
[0048] When generating the three-dimensional representation of each vessel identified in the choroid layer of the eyeball, the computer may generate a vector image representation of each vessel identified in the choroid layer of the eyeball.
[0049] Embodiments may include a non-transitory computer-readable medium storing instructions for choroid analysis of tomography (OCT) imaging. When executed by at least one processor, the instructions may cause the at least one processor to: obtain volumetric data representing three-dimensional imagery of an eyeball based upon a plurality of cross-sectional images received from an OCT imaging device; identify one or more layer boundaries corresponding to a choroid layer of the eyeball based on the plurality of cross-sectional images ofthe volumetric data; identify a plurality of vessel boundaries corresponding to the plurality of vessels of the choroid layer of the eyeball based on the plurality of cross-sectional images of the volumetric data; generate a three-dimensional representation of each vessel identified in the choroid layer of the eyeball; and generate a heatmap based on the three-dimensional representation of each vessel, the heatmap representing variation in diameter of the plurality of vessels the eyeball along centerlines defined by the plurality of vessels.
[0050] The instructions may further cause the at least one processor to, for each vessel in the choroid layer, determine a diameter based upon point cloud data of a point cloud representing the vessel of the eyeball. The heatmap may be generated based upon each diameter determined for each vessel of the plurality of vessels.
[0051] To generate the three-dimensional representation of each vessel in the choroid layer, the instructions may further cause the at least one processor to: for each vessel of the choroid layer of the eyeball : identify a centerline for the vessel as represented by at least one cross-sectional image; determine a plurality of cross-sectional lines extending along the centerline for the vessel; and generate the three-dimensional representation of the vessel as part of a point cloud representing the eyeball based on the centerline and the plurality of cross-sectional lines.
[0052] The instructions further cause the at least one processor to, for each vessel, execute a tensor voting operation for determining the centerline and a centerline point for the vessel.
[0053] The instructions may further cause the at least one processor to: determine a plurality of vessel metrics for the plurality of vessels of the eyeball using the three-dimensional representation of each vessel in the choroid layer, including a diameter for each particular vessel of the plurality of vessels; and generate an output file associated with the heatmap, the output file including the plurality of vessel metrics of the plurality of vessels.
[0054] In some embodiments, a computer-implemented method for choroid analysis of tomography (OCT) imaging is disclosed. The computer-implemented method can include obtaining, by a computer, volumetric data representing three-dimensional imagery of an eyeball based on a plurality of cross-sectional images received from an OCT imaging device. In some implementations, the computer-implemented method can include executing, by the computer, a machine learning model configured to identify one or more layer boundaries corresponding to achoroid layer of the eyeball based on the plurality of cross-sectional images of the volumetric data. In implementations, the computer-implemented method can include determining, by the computer, a subset of the volumetric data representing three-dimensional imagery of the choroid layer in response to identifying the one or more layer boundaries. In at least some implementations, the computer-implemented method can include determining, by the computer, a plurality of centerlines corresponding to a plurality of vessels of the choroid layer based on the subset of volumetric data. The computer-implemented method can include generating, by the computer, a three-dimensional representation of each vessel of the plurality of vessels in the choroid layer corresponding to each centerline of the plurality of centerlines. In implementations, the computer- implemented method can include generating, by the computer, a three-dimensional representation of at least a portion of the choroid layer based on the three-dimensional representation of at least a subset of the vessels.
[0055] In aspects, executing the machine learning model can include providing, by the computer, at least a portion of the volumetric data to the machine learning model to cause the machine learning model to generate an output indicating the one or more layer boundaries corresponding to the choroid layer. Determining the subset of the volumetric data can include determining, by the computer, the subset of the volumetric data based on the indication of the one or more layer boundaries.
[0056] In at least some aspects, executing the machine learning model can include providing, by the computer, at least a portion of the volumetric data to a first encoder of a plurality of encoders along a contracting path of the machine learning model to cause the machine learning model to generate the output.
[0057] In some aspects, determining the plurality of centerlines corresponding to the plurality of vessels can include executing, by the computer, a density-based clustering process to segment the plurality of vessels associated with the subset of volumetric data.
[0058] In aspects, generating the three-dimensional representation of at least a portion of the choroid layer can include generating the three-dimensional representation of at least a portion of the choroid layer, where the three-dimensional representation includes a choroid thickness map.
[0059] In at least some aspects, generating the three-dimensional representation of at least a portion of the choroid layer can include generating the three-dimensional representation of at least a portion of the choroid layer, where the three-dimensional representation includes a choroid vascularity index map.
[0060] In some aspects, the techniques described herein relate to a computer-implemented method, wherein generating the three-dimensional representation of at least a portion of the choroid layer includes: generating, by the computer, a heatmap of at least a portion of a retina of a patient based on the three-dimensional representation of at least a subset of the vessels, the heatmap indicating a thickness of the choroid layer at a plurality of points along the retina of the patient.
[0061] In yet another embodiment, a system is disclosed. The system can include a computing device. The computing device can include at least one processor. The at least one processor can be configured to obtain volumetric data representing three-dimensional imagery of an eyeball based on a plurality of cross-sectional images received from an OCT imaging device. In some implementations, the at least one processor can be configured to execute a machine learning model configured to identify one or more layer boundaries corresponding to a choroid layer of the eyeball based on the plurality of cross-sectional images of the volumetric data. In implementations, the at least one processor can be configured to determine a subset of the volumetric data representing three-dimensional imagery of the choroid layer in response to identifying the one or more layer boundaries; determine a plurality of centerlines corresponding to a plurality of vessels of the choroid layer based on the subset of volumetric data. In some implementations, the at least one processor can be configured to generate a three-dimensional representation of each vessel of the plurality of vessels in the choroid layer corresponding to each centerline of the plurality of centerlines. In at least some implementations, the at least one processor can be configured to generate a three-dimensional representation of at least a portion of the choroid layer based on the three-dimensional representation of at least a subset of the vessels.
[0062] In some aspects, the at least one processor configured to execute the machine learning model can be configured to provide at least a portion of the volumetric data to the machine learning model to cause the machine learning model to generate an output indicating the one or more layer boundaries corresponding to the choroid layer. The at least one processor configuredto determine the subset of the volumetric data can be configured to determine the subset of the volumetric data based on the indication of the one or more layer boundaries.
[0063] In aspects, the at least one processor configured to execute the machine learning model can be configured to provide at least a portion of the volumetric data to a first encoder of a plurality of encoders along a contracting path of the machine learning model to cause the machine learning model to generate the output.
[0064] In some aspects, the at least one processor configured to determine the plurality of centerlines corresponding to the plurality of vessels can be configured to execute a density-based clustering process to segment the plurality of vessels associated with the subset of volumetric data.
[0065] In some aspects, the at least one processor configured to generate the three- dimensional representation of at least a portion of the choroid layer can be configured to generate the three-dimensional representation of at least a portion of the choroid layer, where the three- dimensional representation includes a choroid thickness map.
[0066] In aspects, the at least one processor configured to generate the three-dimensional representation of at least a portion of the choroid layer can be configured to generate the three- dimensional representation of at least a portion of the choroid layer, where the three-dimensional representation includes a choroid vascularity index map.
[0067] In some aspects, the at least one processor configured to generate the three- dimensional representation of at least a portion of the choroid layer can be configured to generate a heatmap of at least a portion of a retina of a patient based on the three-dimensional representation of at least a subset of the vessels, the heatmap indicating a thickness of the choroid layer at a plurality of points along the retina of the patient.
[0068] In yet another embodiment, a non-transitory computer-readable medium is disclosed. The non-transitory computer-readable medium can store instructions thereon that, when executed by one or more processors, cause the one or more processors to obtain volumetric data representing three-dimensional imagery of an eyeball based on a plurality of cross-sectional images received from an OCT imaging device. In some implementations, the instructions can cause the one or more processors to execute a machine learning model configured to identify oneor more layer boundaries corresponding to a choroid layer of the eyeball based on the plurality of cross-sectional images of the volumetric data. In at least some implementations, the instructions can cause the one or more processors to determine a subset of the volumetric data representing three-dimensional imagery of the choroid layer in response to identifying the one or more layer boundaries. In some implementations, the instructions can cause the one or more processors to determine a plurality of centerlines corresponding to a plurality of vessels of the choroid layer based on the subset of volumetric data. In implementations, the instructions can cause the one or more processors to generate a three-dimensional representation of each vessel of the plurality of vessels in the choroid layer corresponding to each centerline of the plurality of centerlines. In at least some implementations, the instructions can cause the one or more processors to generate a three-dimensional representation of at least a portion of the choroid layer based on the three- dimensional representation of at least a subset of the vessels.
[0069] In some aspects, the instructions that cause the one or more processors to execute the machine learning model can cause the one or more processors to provide at least a portion of the volumetric data to the machine learning model to cause the machine learning model to generate an output indicating the one or more layer boundaries corresponding to the choroid layer. The instructions that cause the one or more processors to determine the subset of the volumetric data can cause the one or more processors to determine the subset of the volumetric data based on the indication of the one or more layer boundaries.
[0070] In some aspects, the instructions that cause the one or more processors to execute the machine learning model can cause the one or more processors to providing at least a portion of the volumetric data to a first encoder of a plurality of encoders along a contracting path of the machine learning model to cause the machine learning model to generate the output.
[0071] In aspects, the instructions that cause the one or more processors to determine the plurality of centerlines corresponding to the plurality of vessels can cause the one or more processors to execute a density-based clustering process to segment the plurality of vessels associated with the subset of volumetric data.
[0072] In at least some aspects, the instructions that cause the one or more processors to generate the three-dimensional representation of at least a portion of the choroid layer can causethe one or more processors to generate the three-dimensional representation of at least a portion of the choroid layer, where the three-dimensional representation includes a choroid thickness map.
[0073] In some aspects, the instructions that cause the one or more processors to generate the three-dimensional representation of at least a portion of the choroid layer can cause the one or more processors to generate the three-dimensional representation of at least a portion of the choroid layer, where the three-dimensional representation includes a choroid vascularity index map.
[0074] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory and are intended to provide further explanation of the embodiments described herein.BRIEF DESCRIPTION OF THE DRAWINGS
[0075] The accompanying drawings constitute a part of this specification, illustrate one or more embodiments and, together with the specification, explain the subject matter of the disclosure.
[0076] FIG. 1 is a block diagram of an environment, described in accordance with one or more embodiments herein.
[0077] FIG. 2 is a flow chart illustrating operations of a method for generating predicted images of patient retinas, in accordance with one or more embodiments.
[0078] FIGS. 3A-3H illustrate a non-limiting example of an implementation of techniques for generating predicted images of patient retinas in accordance with one or more embodiments.
[0079] FIGS. 4A-4D illustrate a non-limiting example of an implementation of techniques for generating predicted images of patient retinas at inference time in accordance with one or more embodiments.
[0080] FIG. 5 illustrates a non-limiting example of an image pair with an image and a segmented image, in accordance with one or more embodiments described herein.
[0081] FIG. 6 is a flow chart illustrating operations of a method for detecting instances of disease signatures using a machine-learning architecture, in accordance with one or more embodiments.
[0082] FIGS. 7A-7B show depict dataflow of processes for implementing a segmentation model of a machine-learning architecture, in accordance with one or more embodiments.
[0083] FIGS. 8A-8B depict dataflow of processes for training or implementing a segmentation model of a machine-learning architecture, in accordance with one or more embodiments.
[0084] FIG. 9 is a flowchart illustrating operations of a method for analyzing choroid layer using a machine-learning architecture, in accordance with one or more embodiments.
[0085] FIG. 10 depicts a dataflow of a portion of process for implementing a segmentation model, in accordance with one or more embodiments.
[0086] FIG. 11 depicts a dataflow of a portion of process for implementing a segmentation model programmed and trained to ingest one or more OCT B-scans and output corresponding segmented images for vessel segments, in accordance with one or more embodiments.
[0087] FIG. 12 depicts operations and dataflow of a method of analyzing a choroid layer of a retina, in accordance with one or more embodiments.
[0088] FIG. 13A depicts a model architecture that can be implemented as part of a segmentation model, in accordance with one or more embodiments.
[0089] FIG. 13B depicts a component of the model architecture of FIG. 13A, in accordance with one or more embodiments.
[0090] FIG. 14A depicts operations and dataflow of a method of analyzing a choroid layer of a retina, according to some embodiments.
[0091] FIG. 14B depicts a planar image of a portion of a choroid layer of a retina, according to some embodiments.
[0092] FIG. 14C depicts a choroid thickness map, according to some embodiments.
[0093] FIG. 14D depicts a choroid vascularity indexes map, according to some embodiments.DETAILED DESCRIPTION
[0094] Reference will now be made to the embodiments illustrated in the drawings, and specific language will be used here to describe the same. It will nevertheless be understood that no limitation of the scope of the disclosure is thereby intended. Alterations and further modifications of the features illustrated here, and additional applications of the principles as illustrated here, which would occur to a person skilled in the relevant art and having possession of this disclosure, are to be considered within the scope of the disclosure.
[0095] The conventional techniques and technological solutions available today do not include systems, tools, or platforms that generate predicted images of patient retinas. And by implementing such conventional techniques, when treating adverse conditions (e.g., medical conditions such as retinal detachment, macular degeneration, and / or the like) patients must often participate in lengthy treatment plans given the progression of their conditions. To address the need for techniques that are capable of predicting the early onset of an adverse condition, techniques described herein involve the generation of predicted images of patient retinas. More specifically, embodiments described herein include a system for generating predicted images of patient retinas, the system comprising: at least one processor programmed to: receive data associated with a first image, the first image captured by an optical coherence tomography (OCT) device while imaging a retina of a patient; generate a segmented image based on the first image, the segmented image associated with at least one boundary line denoting a boundary between a first layer of the retina and a second layer of the retina; determine an image pair based on the first image and the segmented image; and execute a model using the image pair to cause the model to generate an output representing a predicted image of the retina of the patient
[0096] By implementing the techniques described herein, systems (e.g., computing devices) can receive data from one or more sources (e.g., OCT devices) and generate predicted images of patient retinas at a point in time in the future. These predicted images may represent transformations in the eyeballs of patients that can be interpreted (e.g., by devices and / or by clinicians) as indicating future adverse conditions. And once such conditions are identified, clinicians can take steps proactively to prevent or treat such adverse conditions. By preventing or treating these conditions, patient health outcomes may be improved and the chances that patients will later need surgery or other interventions may be significantly reduced if not eliminated.
[0097] Further, by virtue of implementing the techniques and systems described herein, symptoms that are otherwise unidentifiable and / or uninterpretable by humans may be identified. For example, in some cases where humans are unable to perceive visual differences between retinal layers, systems described herein may be able to perceive such visual differences and determine that one or more adverse conditions are advancing from extremely early stages. This can enable clinicians to not only treat such adverse conditions early, but also determine when precisely to start such treatment so as to optimize the patient’s health outcome.
[0098] Another shortcoming in current technologies in supporting the health of a patient’s eyeball includes a capability to detect instances of certain disease signatures and age-related macular degeneration. Drusen is a key clinical marker that can aid in dry, age-related macular degeneration prognosis, but due to the quantity of data and observations required in practice and the variability of drusen-related data, tracking and detecting drusen is difficult over time. While manual segmentation and analysis is possible to derive and identify the biomarkers of drusen, such approaches are not used in common practice because the segmentation and analysis is a timeconsuming and labor-intensive task.
[0099] Embodiments described herein include a computing system executing software functions of an ophthalmic image processing tool to detect the presence of disease signatures and age-related macular degeneration. For instance, the computing system may detect and segment instances of drusen in a diverse patient population receiving ophthalmic care. The computing system executes software implementing a machine-learning architecture trained to identify OCT images that contain drusenoid lesions and mark those lesions. During training, the machinelearning models of the machine-learning architecture may be trained to be robust against image artifacts and noise. The computing system may calculate biomarkers based on segmented regions of the OCT that can be used for the prediction and evaluation of disease progression and treatment response in age-related macular degeneration.
[0100] In some embodiments, a computing system for disease detection may perform operations for drusen identification and segmentation using a machine-learning model trained for analyzing layers and attributes of en-face OCT images. The OCT images may be referenced and analyzed for identifying instances of age-related macular degeneration as the OCT images can indicate various forms of drusen or other forms. However, traditional B-scan OCT images arelimited by their cross-sectional view of the retina, which prevents analysis of the whole retina in a single image. These aspects are crucial for the evaluation of drusen, which are characterized by a multitude of lesions that span the macular region. En-face OCT views overcome this hurdle by providing physicians with a view of the entire retina while retaining the usefulness of OCT imaging technology.
[0101] In some embodiments, a computing system executes software functions of a multimodal ophthalmic image quantification tool having operations for implementing machinelearning models and image processing to analyze aspects of a CL in OCT images. The computing generates or otherwise facilitates biomarker quantification for various downstream operations of disease screening tools. Generating these biomarker quantification outputs may include, for example, using the OCT scans to generate CL thickness maps, CL vascularity index (CVI) maps, CL inner / outer surface contour maps, CL area / volume measurements, CL vessel volume measurements, CL vessel diameter measurements, and CL vessel diameter heatmaps. In some cases, generating these biomarker quantification outputs may include generating area / volume measurements of various retinal lesions including subretinal fluid (SRF), pigment epithelium detachment (PED), and intra-retinal fluid using OCT scans. In some cases, generating these biomarker quantification outputs may include generating measurements of retinal lesions including hard exudates using color fundus (CF) photographs. Additionally or alternatively, generating these biomarker quantification outputs may include executing software programming of a multimodal manual measurement tool for measuring thickness, area, and volume of lesions, as well as various layers of the posterior segment of the eye, among other types of outputs. The executable software may perform multimodal image registration including, for example, registration across various 2D retinal image modalities; (ii) registration of longitudinal 2D retinal images obtained from the same modality; (iii) volumetric registration of longitudinal volumes; (iv) registration of 2D image modalities such as CF, fundus autofluorescence (FAF) with 3D OCT volumes; and (v) registration of 2D and 3D retinal images with Humphrey visual field data.
[0102] The quantification outputs from analyzing the CL may be beneficial for screening and management of choroidal and retinal diseases, including age-related macular degeneration (AMD), central serous chorioretinopathy (CSCR), diabetic retinopathy (DR), macular edema (ME), and glaucoma, among others. By generating and providing the quantification of biomarkers(e g., CL thickness, CVI, retinal lesion measurements), the software of the quantification tool may facilitate early detection and accurate monitoring of disease progression. Clinicians, for example, may interact with a graphical user interface of the quantification to assess disease severity, track changes over time, and tailor treatment strategies for optimal patient outcomes. Additionally, the software’s image registration capabilities may enhance data integration and analysis of multimodal imaging data, offering a comprehensive approach to disease assessment and management.
[0103] FIG. l is a block diagram of an environment 100, described in accordance with one or more embodiments herein. The environment 100 may include OCT device 110a, OCT device 110b, and OCT device 110c (referred to collectively as OCT devices 110 and individually as OCT device 110, unless otherwise specified), client device 130a, client device 130b, and client device 130c (referred to collectively as client devices 130 and individually as client device 130, unless otherwise specified), a server 140, and / or a database 150. The OCT devices 110, client devices 130, server 140, and database 150 may all interconnect (e.g., establish a connection to communicate with one another) via a network 120. The network 120 may be a wide area network (such as the Internet), a local area network (LAN), or any other kind of network.
[0104] The OCT devices 110 include one or more devices (including one or more computing device) comprising hardware and software components capable of performing the various processes described herein. For example, the OCT devices 110 may be any device including a memory and a processor capable of communicating, via the network 120 with one or more other devices of FIG. 1. In some implementations, the OCT devices 110 include one or more of a light source, an interferometer, a sample arm, a reference arm, a detector, a signal processor, a scanner, a processor, memory, and a display. The OCT devices 110 are configured to be in communication with the client devices 130, the server 140 and the database 150 via network 120. The OCT devices 110 provides (e.g., transmits) data to the client devices 130 and the server 140 as described herein. While the server 140 is described as managing communication between the OCT devices 110 and the database 150, it will be understood that the OCT devices 110 may communicate directly with the database 150 to upload data to the database 150 to be stored. In some embodiments, the OCT devices 110 are associated with clinicians as described herein.
[0105] The client devices 130 include any computing device comprising hardware and software components capable of performing the various processes described herein. For example,the client devices 130 may be any device including a memory and a processor capable of communicating, via the network 120 with one or more other devices of FIG. 1. Non-limiting examples of the client devices 130 include desktop computers, mobile devices (e.g., cellular phones and tablets), and / or the like. The client devices 130 are configured to be in communication with OCT devices 110, server 140, and database 150 via network 120. In some embodiments, the client devices 130 are associated with clinicians and / or patients as described herein.
[0106] The server 140 includes any computing device comprising hardware and software components capable of performing the various processes described herein. For example, the server 140 may be any device including a memory and processor capable of communicating, via the network 120 with one or more other devices of FIG. 1. Non-limiting examples of the server 140 includes data centers, server computers, workstation computers and / or the like. The server 140 is configured to be in communication with the OCT devices 110 and the client devices 130 via network 120. In some embodiments, the server 140 is associated with one or more clinicians and / or one or more healthcare providers (e.g., hospitals, public and / or private research organizations, and / or the like).
[0107] The database 150 includes any computing device comprising hardware and software components capable of performing the various processes described herein. For example, the database 150 may be any device including a memory and processor capable of communicating, via the network 120 with one or more other devices of FIG. 1. Non-limiting examples of the database 150 includes data centers, server computers, workstation computers and / or the like. The database 150 is configured to be in communication with the OCT devices 110, the client devices 130, and the server 140 via network 120. The database 150 may obtain (e.g., receive) data from the OCT devices 110, the client devices 130, and / or the server 140. In some embodiments, the database 150 is associated with (e.g., under the control of) the client devices 130, the server 140 and / or one or more clinicians as described herein.
[0108] The data described herein may include image data associated with one or more images of a portion of an eyeball of a patient. For example, the image data may include data generated by OCT devices 110 while imaging the eyeballs (e.g., the posterior portion of the eyeballs) of patients. Images may be generated by various OCT devices 110 at instant points in time or at points in time over a period of time. And the images may be correlated with one or morepatients, one or more periods of time, and / or the like. The database 150 may contain data related to users or research projects (e.g., patient data, researcher data, clinician data, research data), which may be stored in database records (e.g., patient records, research records). Non-limiting examples of the types of data in the database records may include a patient identifier (ID), research ID, timestamps of data inputs, laterality (e g., vitals, imaging, social determinants of health), data modality (e.g., CF, OCT, FAF), operating mode (e.g., single scan, volume scan), raw image data, and image analysis data, among others.
[0109] FIG. 2 is a flow chart illustrating operations of a method 200 for generating predicted images of patient retinas, in accordance with one or more embodiments. In some implementations, one or more of the functions described with respect to method 200 may be performed (e.g., completely, partially, and / or the like) by a server that is the same as, or similar to, server 140. In some implementations, one or more of the functions described with respect to method 200 may be performed (e.g., completely, partially, and / or the like) by another device or group of devices separate from and / or including the server, such as by one or more OCT devices that are the same as, or similar to, OCT devices 110 of FIG. 1, one or more client devices that are the same as, or similar to, the client devices 130 of FIG. 1, and / or a database that is the same as, or similar to, the database 150 of FIG. 1.
[0110] At operation 210, the method 200 includes receiving, by a server, data associated with a first image and a second image. In some implementations, the first image and the second image may include an image of at least a portion of an eyeball of a patient. For example, the server may receive data associated with a first image, where the first image is represented as a B-scan. The B-scan may represent a two-dimensional cross-section of a retina of a patient (see, e g., the image generated by OCT device 410 of FIG. 4). For example, the B-scan may represent a two- dimensional cross-section of an anterior or posterior of the retina of the patient. The B-scan may include a Cirrus OCT B-scan, a standard OCT B-scan, or an enhanced depth imaging OCT B-scan (sometimes referred to as an “EDI OCT” image, a more detailed scan when compared to the standard OCT image that enables visualization of the sclera of the eyeball).[0U1] In some implementations, the first image and the second image may represent multiple layers associated with (e.g., in proximity to) the retina of the patient. For example, the first image and the second image may represent the retina (e.g., one or more layers of the retina),the choroid, and / or the sclera of the patient. In some implementations, the one or more layers of the retina represented include the retinal nerve fiber layer (RNFL), the ganglion cell layer (GCL), the inner plexiform layer (IPL), the inner nuclear layer (INL), the outer plexiform layer (OPL), the outer nuclear layer (ONL), the inner segment / outer segment junction (IS / OS), the retinal pigment epithelium (RPE). In some implementations, the first image also represents the vitreous chamber of the eyeball of the patient.
[0112] In some embodiments, the server receives data associated with the first image and the second image, where the first image and the second image include OCT volume scans. In this example, the OCT volume scans are represented as a plurality of first images or second images (e.g., a plurality of B-Scans such as a plurality of standard OCT B-scans, a plurality of EDI OCT B-scans, a plurality of Cirrus OCT B-scans, and / or the like representing an area of the eyeball of the patient.
[0113] In some implementations, the server may receive the data associated with the first image from one or more OCT devices. For example, the server may receive the data associated with the first image from the one or more OCT devices based on the one or more OCT devices imaging eyeballs of patients (e.g., generating Cirrus OCT B-scans, standard OCT B-scans, EDI OCT B-scans, OCT volume scans, and / or the like of eyeballs of patients). In some implementations, the one or more images are each associated with a point in time. For example, the one or more images generated by the OCT devices while imaging the eyeballs of patients may be associated with a timestamp. The timestamp may indicate the date and time at which the OCT devices performed the imaging of the eyeballs of the patients.
[0114] The server may receive the data associated with the first image where the first image is included in a set of images including images corresponding to multiple patients. In some implementations, the set of images includes other images captured by the OCT devices (e.g., other B-scans), where the set of images include images of multiple patients’ eyeballs. In such examples, the server may associate each image with a point in time by including a timestamp with each image. As discussed below, the point in time may be chronologically before or after a point in time at which one or more other images of the corresponding patients’ eyeballs were generated by the OCT devices.
[0115] The server may associate the images generated by the OCT devices with a corresponding patient. For example, as the eyeball of a patient is imaged over time (e.g., at multiple points in time corresponding to different visits to the clinician) the data associated with each image may be provided (e.g., transmitted) to the server. The server may then store the data associated with the images of the eyeball of the patient in association with other data that is previously- generated in a database (e.g., a database that is the same as, or similar to, the database 150 of FIG. 1). In this way, the server may index multiple images of patients’ eyeballs to later display to clinicians or when to use when training models as described herein.
[0116] When generating the images of the patients’ eyeballs, the OCT devices may generate images consecutively starting at the point in time at which the first image is generated. For example, to generate an OCT volume scan (e.g., a three-dimensional image of at least a portion of the eyeball of the patient) the OCT devices may generate multiple B-scans and align (e.g., register) the B-scans with one another. The position of the multiple B-scans may be registered with one another or to one individual B-scan as a reference to generate the OCT volume scan (the three- dimensional representation) of the eyeball of the patient. When registering against a reference, the server may align the multiple B-scans of an OCT volume scan relative to a B-scan that is transverse to the multiple B-scans (referred to as an en-face image) that are being aligned. In some implementations, the OCT volume scan may be generated at the same time the one or more individual B-scans are generated. Examples of multiple B-scans registered relative to one another are represented by FIG. 3F.
[0117] As described above, the server may receive data associated with a second image. For example, the server may receive data associated with a second image captured by one or more OCT devices. The second image may be an image that is similar to the first image of the eyeball of a patient that is captured at a second point in time. For example, the patient may visit the clinician multiple times (e.g., over weeks, months, and / or years) and the clinician may operate the OCT device during one or more visits to generate the images of the portion of the eyeball of the patient. In some implementations, the first image and the second image may correspond to a portion of interest of the eyeball of the patient. For example, the image may be associated with one or more specific structures of the eyeball (e.g., the optic nerve head). In this example, where the first and second image correspond to the specific structures, the images may be captured todocument a chronological progression of updates to the specific structures of the eyeball of the patient.
[0118] The second image may be captured at a point in time relative to a point in time that the first image is captured. For example, the second image may be captured at a point in time that is chronologically later than the point in time at which the first image was captured. In some examples, the period of time between the first point in time and the second point of time is a predetermined period of time. The predetermined period can include a week, a month, a year, two years, and / or the like. In some implementations, the server may receive other images besides the first image and the second image at different points in time. For example, the server may receive a third image, etc., at points in time that are chronologically later and represent similar periods as the period from the first point in time to the second point in time. In this way, the server may periodically receive image data from one or more OCT devices, the image data being associated with one or more particular patients. The server may then store the image data in the database for recall at a later point in time to be analyzed as described herein.
[0119] In some embodiments, the server may calculate a histogram based on the first image, the second image, and / or one or more other images. For example, the server may calculate a histogram for the first image, the second image, and / or the other images stored in the database. The histograms may represent a graphical representation of the distribution of pixel intensities in the one or more images. In some implementations, the server calculates an average histogram based on the first image, the second image, and / or the one or more other images. For example, the server may calculate an average histogram when generating a dataset. In some implementations, the server may update the one or more images described herein based on the average histogram and / or the histogram of each individual image. For example, the server may update the one or more images described by matching the histogram of each individual image to a certain histogram. The certain histogram may include one or more other images and / or the average histogram of the images stored in the database. In these examples, the server may update the values of a given image or the server may generate a dataset based on updating the one or more images. In this way, the server may normalize the images captured by the OCT devices. This may be important in the case where certain OCT devices are calibrated differently from other OCT devices and generate imageshaving different profiles (e.g., histograms) or when multiple image types (e.g., Cirrus OCT scans and EDI OCT scans) are captured for a given patient.
[0120] At operation 220, the method 200 includes generating, by the server, a segmented image based on the first image and the second image. For example, the server may generate the segmented image based on the first image and the second image (see, e.g., segmented image 420 of FIG. 4) where the first image and the second image are represented as an OCT B-scan or an OCT volume scan. In some implementations, the server generates the segmented image based on the first image, the second image, and / or one or more other images. For example, as images are generated by OCT devices and stored in the database, the server may update the images stored therein by segmenting the images. In this way, the server may generate a dataset to be used to train a model to predict images of patient retinas (discussed with respect to operation 240 and FIGS. 3G and 3H).
[0121] In some embodiments, the server generates the segmented images based on (e.g., using) a model. The model may be the same as, or similar to, the model described with respect to operation 240. For example, in the case of a GAN, the server may provide images (e.g., nonsegmented images) to the generator network of the GAN and the server may provide corresponding target images (e.g., segmented images) to the discriminator network of the GAN. The server may then cause the model to be updated (e.g., by updating one or more weights associated with the generator network or the discriminator network) and repeat this process until the GAN converges. The server may perform this process prior to training the model on image pairs of images captured at different points of time as described herein. In embodiments, the server may repeat this process for multiple types of images (e.g., EDI OCT B-scans, Cirrus OCT B-scans, and / or the like) and corresponding segmented images (e.g., segmented EDI OCT B-scans, segmented Cirrus OCT B- Scans, and / or the like). The server may repeat this process until the model converges for each image type. In this way, the server may train the model to generate segmented images for multiple image types.
[0122] When pretraining and / or conditioning the GAN, the server may provide data associated with an image pair (e.g., a B-scan and a segmented B-scan) to the generative model of the GAN, where the image pair is based on an input image that was updated. For example, where the input image is updated to match an average histogram, the input image may then be segmentedsuch that the image pair includes the input image that is matched to the average histogram and a corresponding segmented image that is similarly matched to the average histogram. The image pair may then be provided to a GAN which includes the following encoder layers: a first encoder layer with 64 convolutional filters, a second encoder layer with 128 convolutional filters, a third encoder layer with 256 convolutional filters, a fourth encoder layer with four layers of 512 convolutional filters; and corresponding decoder layers: a fourth decoder layer of four layers of 512 convolutional filters, a third decoder layer with 256 convolutional filters, a second decoder layer with 128 convolutional filters, and a first decoder layer with 64 convolutional filters. The output layer of the GAN may be configured to provide data associated with a segmented image as an output. The output image may be a segmented image including one or more boundary lines, as described below.
[0123] The GAN may also include a discriminative model (e.g., a PatchGAN model and / or the like). The server may provide the output of the generative model of the GAN to the discriminative model. Additionally, or alternatively, the server may provide the corresponding segmented image (e.g., a target image) to the discriminative model to cause the discriminative model to generate the output. The discriminative model may be configured to generate an output indicating whether the output of the generative model of the GAN is a real image or a not real (e.g., fake) image. In some implementations, the discriminative model includes a first convolutional layer of 64 convolutional filters, a second convolutional layer of 128 convolutional filters, a third convolutional layer of 256 convolutional filters, and a fourth convolutional layer of 512 convolutional filters. The output may be an indication that is binary (e.g., real or not real) or may be a value representing a likelihood that an image is real or not real. An example flowchart is included below at FIG. 5 representing the implementation involving training of a generator network and a discriminator network as described.
[0124] The server may train the model (e.g., pretrain and / or condition the GAN) through multiple training sessions. For example, the server may first take image pairs of standard OCT images and their corresponding segmented images (referred to in combination as “standard OCT image pairs”) and train the model based on the standard OCT image pairs. The server may then take images pairs of EDI OCT images and their corresponding segmented images (referred to in combination as “enhanced OCT image pairs”) and train the model based on the enhanced OCTimage pairs. In some implementations, the standard OCT image pairs or the enhanced OCT image pairs include pairs of images that were previously matched to an average histogram. In some embodiments, the standard OCT image pairs and the enhanced OCT image pairs may then be used to train the model prior to being matched to the average histogram. In this example, the model may then be retrained using the standard OCT image pairs and the enhanced OCT image pairs based on the standard OCT image pairs and the enhanced OCT image pairs being matched to the average histogram.
[0125] The server may generate the segmented images, where the segmented images are associated with (e.g., represent) at least one boundary line. For example, in the case of an OCT image such as the first image, the second image, and / or any other suitable image generated by the OCT devices, the server may generate the segmented image such that the boundaries between the structures of the eyeballs are denoted. In one example, the server generates the segmented image such that certain layers of the retina (described herein) are separated from one another. The segmented images may then be reviewed (e.g., by an individual) to confirm that the segmentations (e.g., the differentiated layers of the retina) are correct. For example, the server may communicate with a client device (e.g., a client device that is the same as, or similar to, client devices 130) and transmit data associated with the one or more segmented images to the client devices. The client devices may then display (e.g., via a display device of the client devices, not explicitly shown in FIG. 1) the segmented images and a clinician may provide input confirming that the segmentation is appropriate. In the case where the segmentation is not appropriate, the clinician may provide an indication via the corresponding client device that the segmentation is not appropriate and / or an indication regarding how to update the segmented image to correct the segmentation. In this case, the clinician may provide input via the client device that causes one or more points along the segmentation lines to move in accordance with the boundary between the different structures of the eyeball of the patient. The server may then update the segmented image and store the segmented image in the database. In some implementations, the server may then store the segmented images in the database.
[0126] In some embodiments, the server may initially generate the target images (the target segmented images) based on one or more segmentation algorithms. For example, the server may generate the target images based on one or more threshold-based techniques. In these examples,the server may implement thresholding that is based on pixel intensity values by comparing adjacent pixels and determining whether the adjacent pixels satisfy a threshold difference in intensity values. The server may also implement adaptive thresholding whereby the threshold difference is determined (e.g., adjusted) based on one or more values within a predetermined distance from a given pixel. In some implementations, the server generates the segmented images based on one or more graph-based techniques such as by fitting active contours to the images. And in some examples, the server may smooth the one or more segmentation lines.
[0127] At operation 230, the method 200 includes determining, by the server, an image pair. For example, the server may determine one or more image pairs based on corresponding pairs of unsegmented images (e.g., the first image or the second image) and segmented images (e.g., the segmented image corresponding to the first image and the segmented image corresponding to the second image, respectively). The server may then store the image pair including the unsegmented first image and the unsegmented second image in the database.
[0128] In some embodiments, the server may determine one or more image pairs based on one or more images of the same portion of the eyeball. For example, the server may receive a B- scan of a portion of an eyeball of a patient (e.g., the first image) and the server may receive another B-scan of the portion of the eyeball of the patient (e.g., the second image). In this example, the server may determine a correspondence between the first image and the second image (as described below), and the server may store the images in correspondence with one another in the database. In this way, the server may index images based on a chronological progression, whereby each image in the index represents the portion of the eyeball of the patient at different points in time (e.g., ever year, every other year, and / or other time period).
[0129] The server may determine a correspondence between one or more portions of the one or more images. For example, the server may determine a correspondence between one or more portions of one or more images in an image pair that represent the same portion of an eyeball, where the image pair includes a first image captured at a first point in time and a second image captured at a second point in time. In some examples, the server may determine the correspondence between the one or more portions of the one or more images based on the server determining that one or more layers represented by corresponding segmented images are associated with one another (e.g., match). The one or more layers may represent one or more layers of a structure (e.g.,a retina) of the eyeball of a patient. When determining the correspondence between the one or more portions of the one or more images, the server may compare the one or more portions of the images (representing one or more layers of the eyeball of the patient) and determine that the images correspond to one another. When comparing the one or more portions, the server may compare the one or more boundary lines included in the one or more images. In this way, the server may determine a correspondence across images even in cases where one or more boundary lines do not clearly correspond with one another (e.g., in the case where an adverse condition such as macular degeneration has caused one layer to move and / or not be present in the same position over time).
[0130] In some embodiments, the server may determine a correspondence between a first image captured at a first point in time and a second image captured at a later point in time (e.g., a point in time one to two years later). In this way, the server may further determine a correspondence between multiple initial images over periods of time to be used during training as described with respect to operation 240, below. In some embodiments, the server may determine the correspondence between the first image and the second image based on registration of the first image and the second image. For example, the server may extract en-face images associated with the first image and the second image (where the first image and the second image are associated with OCT volume scans). Where the extracted en-face images correspond to scans at different points in time (e.g., a first visit by a patient and a second visit by a patient), the server may compare the en-face images and register the en-face images that correspond with one another. In some embodiments, the server may then determine a correspondence between the first image and the second image relative to the corresponding en-face images. For example, where the first image and the second image are included in corresponding OCT volume scans, the server may determine a correspondence between the corresponding OCT volume scans based on the position of the en- face images relative to the OCT volume scans. In some embodiments, the server may transform one OCT volume scan based on the registration of the corresponding en-face images. In this way, the server may account for structural changes when aligning multiple B-scans represented by OCT volume scans at different points in time.
[0131] In some embodiments, the server may pre-process one or more images. For example, the server may pre-process one or more images stored in the database. In this example, the server may preprocess the one or more images by cropping the one or more images (e.g., from1024x512 pixels to 900x512 pixels and / or to 512x512 pixels). In examples, the server may pre- process the one or more images before determining the average histogram for the images stored in the database. In some implementations, the server may pre-process the one or more images stored in the database based on (e.g., prior to or after) the server determining the one or more boundaries represented in the one or more images.
[0132] At operation 240, the method 200 includes executing, by the server, a model using the image pair to cause the model to generate an output. For example, the server may execute the model using the image pair by providing (e.g., transmitting and / or making available for download) the image pair to a model as input to the model to cause the model to generate an output. During training, weights of the model may be updated based on the model generating the output. The weights may be updated iteratively until the output of the model satisfies a threshold difference. For example, the server may iteratively update the weights of the model and provide the image pair to the model to cause the model to generate an output representing one or more images and / or one or more features of the one or more images of the image pair. In this example, the weights may be updated based on the server calculating a cross entropy between the boundaries associated with the predicted image and boundaries of the segmented image of an image pair used to train the model.
[0133] The server may provide the image pair to a model to cause the model to generate an output, where the model involved is a generative adversarial network (GAN). For example, the GAN may be associated with a U-net architecture (e.g., a Residual U-net and / or the like). In this example, the GAN can include a U-net architecture that includes a series of encoders and decoders, the encoders associated with a contracting path (or an encoder path) and the decoders associated with an expanding path (or a decoder path). A first encoder may receive data associated with the image pair, perform one or more functions, and provide an output to a second encoder along the contracting path as well as an output to a corresponding decoder along the expanding path.
[0134] The output of the model (e.g., a GAN and / or the like) may include data associated with a predicted image. For example, the output of the model can include data associated with a predicted image of at least a portion of a retina of a patient. The predicted image may visually represent one or more changes to at least a portion of the eyeball of the patient that are expected to occur over a period of time. In some implementations, the output includes data associated witha predicted image, where the predicted image is associated with a predetermined period of time. For example, where the model is trained using a set of image pairs, each image pair including an image captured at a first point in time and an image captured at a second point in time, respectively, the model may be trained to generate predicted images at a point in time in the future. In this example, the point in time in the future may correspond to the period of time between the point in time that the first image was captured and the point in time that the second image was captured. Additionally, or alternatively, the model may be trained based on image pairs captured at varying points in time and the model may provide as an output data associated with images irrespective of the amount of time between when the images were captured.
[0135] The server may provide the image pair including a first image and a corresponding segmented image to the GAN. For example, the server may provide the image pair to the generator network of the GAN to pre-train or condition the GAN to predict images and their corresponding segmentations. Additionally, or alternatively the server may provide an image pair including the first image and a corresponding second image to the GAN. In this example, the first image and the second image may be B-scans of the same portion of an eyeball of a patient. In this way, the GAN may be conditioned to receive B-scans of an eyeball of a patient and generate predicted images of a patient at a period of time in the future.
[0136] The server may cause a device to display one or more images. For example, the server may cause one or more client devices to display one or more images based on the server generating one or more predicted images. Additionally, or alternatively, the server may cause the one or more client devices to display the one or more images stored in the database. For example, the server may receive a request to display one or more images stored in the database (e.g., one or more OCT images of a patient captured at one or more points in time) and the server may retrieve the images stored in the database. The server may then generate the data to cause the client device that requested the images to display the images retrieved from the database. In examples, the server may transmit data associated with the one or more predicted images where the data is configured to cause the predicted image to be presented on a display device (e.g., a screen) of the client device.
[0137] The server may cause one or more additional images to be stored in the database. For example, the server may cause one or more fundus autofluorescence (FAF) images, one or more color fundus (CF) images, and / or one or more registered FAF-CF images (e.g., images wherean FAF is positioned relative to structures represented in a CF image as an overlay or underlay to the CF image). In some implementations, the one or more FAF and / or CF images may be registered relative to one or more B-scans described herein. In such an example, the one or more B-scans may represent an image of the structure of the eyeball of the patient that is positioned transverse to the structures similarly represented in the FAF and / or CF image. In some implementations, the client device may transmit a request to the server for the FAF images, CF images, and / or FAF-CF images. For example, the client device may request the images based on (e.g., before, during, and / or after) a visit between a patient and a clinician. In such examples, the server may retrieve the requested images from the database and transmit the images to the client device. This, in turn, may cause the client device to display the requested image on a display device associated with the client device.
[0138] In some embodiments, the server may store data associated with the images described herein in a database. For example, the server may store the images described herein in association with data related to users or research projects (e.g., patient data, researcher data, clinician data, research data), which may be stored in database records (e.g., patient records, research records). Non-limiting examples of the types of data in the database records may include a patient identifier (ID), research ID, timestamps of data inputs, laterality (e.g., vitals, imaging, social determinants of health), data modality (e.g., CF, OCT, FAF), operating mode (e.g., single scan, volume scan), raw image data, and image analysis data, among others. In some embodiments, the server may store data associated one or more of a visit date, an image modality, a scan type, an age or age range associated with patients corresponding to one or more images, genders associated with one or more patients corresponding to one or more images, an international classification of disease (ICD) code associated with the one or more patients, and / or the like. Once stored, the server may retrieve some or all of the data described herein and cause the data to be displayed via a display device. In some embodiments, the data may be stored as objects in the database in correspondence with one or more fields, the one or more fields corresponding to the data described above.
[0139] One or more of the functions described herein may be associated with a Pythonbased script. For example, one or more of the functions described herein associated with the preprocessing of the images and / or training of models described herein may be associated withPython-based scripts. It will be understood that other suitable tools may be used by one of ordinary skill in the art to implement the functions described herein.
[0140] FIGS. 3A-3H illustrate a non-limiting example of an implementation 300 of techniques for generating predicted images of patient retinas in accordance with one or more embodiments. In some implementations, one or more of the computing devices described may be the same as, or similar to, one or more of the computing devices of FIG. 1.
[0141] As illustrated by an operation 360, an OCT device 310 generates OCT B-scans and OCT volume scans of the eyeball of one or more patients. In some implementations, the OCT device 310 is the same as, or similar to, the OCT devices 110 of FIG. 1. In some embodiments, the OCT B-scans may include Cirrus OCT B-scans, EDI OCT B-scans, OCT volume scans, OCT B-scans (captured transverse to the OCT B-scans of a volume scan and referred to as en-face images), and / or the like.
[0142] As illustrated by an operation 362, the OCT device 310 transmits data associated with the images (the OCT B-scans and OCT volume scans) to a server 340. In some implementations, the server is the same as, or similar to, the server 140 of FIG. 1.
[0143] In some implementations, the server 340 may map one or more of the images to one or more other images. For example, the server 340 may map one or more EDI OCT B-scans to one or more Cirrus OCT B-scans. When doing so, the server 340, may calculate a histogram for all of the Cirrus OCT B-scans, calculate an average of the histograms, and for each EDI OCT B- scan, map the histogram of the EDI OCT B-scan to the average of the histograms. The server 340 may then update (e.g., add or replace) one or more images received from the OCT device 310 with one or more of the mapped EDI OCT B-scans.
[0144] As illustrated by operation 364, the server 340 generates segmented images. For example, the server 340 may generate the segmented images based on the data associated with the images received from the OCT device 310 and / or the updated images generated when mapping the EDI OCT B-scans to the average of the histograms. In some implementations, the server 340 may generate the segmented images using one or more segmentation techniques and / or using a model 342.
[0145] As illustrated by operation 366, the server 340 transmits data associated with images (e.g., B-scans and corresponding segmented B-scans) to a client device 330. In some implementations, the client device 330 is the same as, or similar to, the client devices 130 of FIG. 1. In some implementations, the server 340 transmits data associated with images to the client device 330 to cause the client device 330 to display the images via a display device to a user (e.g., a clinician). The user may then provide input to the client device 330 indicating whether the boundary lines represented in the segmented images are accurate.
[0146] As illustrated by operation 368, the server 340 receives a response from the client device 330. In some implementations, the response may include an indication associated with whether the boundary lines represented by the segmented images are accurate.
[0147] As illustrated by operation 370, the server 340 trains a model 342 based on pairs of Cirrus OCT B-scans and segmented images (target images). In some implementations, the model 342 is a GAN as described herein (e.g., a Pix2Pix GAN and / or the like).
[0148] During training, the model 342 may be trained as follows. For a number of training steps (e.g., 40,000 training steps), provide a Cirrus OCT image and a segmented EDI OCT image to a first layer of a GAN. The GAN may include a generator network configured to receive a histogram mapped image (e.g., a 512x512 image), analyze the image, and provide outputs to successive layers of the GAN as described herein. In some embodiments, the generator network of the GAN may be configured to output an image (e.g., a 512x512 image) having one or more boundary lines. In some embodiments, the GAN may include a discriminator network. The discriminator network of the GAN may be configured to receive the image output by the generator network and a boundary image (e.g., a target image representing one or more boundary lines). The discriminator network may be configured to output a determination about whether the image that was output by the generator network is real or fake.
[0149] As illustrated by operation 372, the model 342 may be retrained. In some embodiments, the model 342 may perform the following steps. For a number of training steps (e.g., 40,000 training steps), provide an EDI OCT image and a segmented EDI OCT image to a first layer of a GAN. The GAN may include a generator network configured to receive a histogram mapped image (e.g., a 512x512 image), analyze the image, and provide outputs to successivelayers of the GAN as described herein. In some embodiments, the generator network of the GAN may be configured to output an image (e.g., a 512x512 image) having one or more boundary lines. In some embodiments, the GAN may include a discriminator network. The discriminator network of the GAN may be configured to receive the image output by the generator network and a boundary image (e.g., a target image representing one or more boundary lines). The discriminator network may be configured to output a determination about whether the image that was output by the generator network is real or fake.
[0150] In some implementations, the server 340 may crop the images before training the model 324. For example, the server 340 may crop the images from a first size (e.g., a size between 1024x512 and 900x512) to a second size (e.g., 512x512). In this way, the server 340 may ensure that the inputs and outputs to the model 342 are normalized before training the model 342. During training, the server 340 may calculate the cross entropy (loss) between the predicted boundaries and the boundaries in the target image and update one or more weights of the model. This process may be repeated until the model converges.
[0151] In some embodiments, the server 340 may smooth one or more of the boundary lines. For example, the server 340 may smooth one or more of the boundary lines. For example, the server 340 may smooth the one or more boundary lines determined by the server 340 that are included in one or more segmented images. In examples, the server 340 may smooth the one or more boundary lines determined by the server 340 that are included in one or more segmented images using logical regression techniques that fit smooth curves through data points by forming weighted regressions in localized portions of the segmented images. The server 340 may then train one or more models as described here in based on the server 340 smoothing the one or more boundary lines.
[0152] As illustrated by operation 374, the server 340 causes the model 342 to segment images. For example, the server 340 may cause the model 342 to segment some or all of the images received from the OCT device 310. In this way, the server 340 may prepare to determine one or more correspondences between one or more images as described herein.
[0153] As illustrated by operation 376, the server 340 registers en-face images. For example, the server 340 may extract one or more en-face images corresponding to one or moreOCT volume scans. The server 340 may then register the en-face images extracted with the corresponding OCT volume scan.
[0154] As illustrated by operation 378, the server 340 registers the OCT B-scan images based on corresponding en-face images. For example, the server 340 may determine a position of the en-face images relative to the corresponding OCT volume scans (referred to as registration). In this way the server may determine the relative position of all of the OCT B-scan images associated with an OCT volume scan.
[0155] When registering the OCT B-scan images, server 340 may segment each layer of each image of an OCT volume scan, obtain enface images from the OCT volume image (where the OCT volume image includes 40 samples) and obtain a transformation to register both enface images (e g., of multiple OCT volume scans). In some implementations, the server 340 may implement techniques for aligning greyscale images (e.g., en-face images) based on pixel intensities. This process may be repeated to align each en-face image with corresponding en-face images (e.g., across multiple OCT volume scans) to determine which en-face images correspond to other en-face images. Based on the registration of the en-face images, the server 340 repeat the alignment process to align each OCT B-space image with corresponding OCT B-space images (e.g., across multiple OCT volume scans) to determine which OCT B-space images correspond to other OCT B-space images. Similar to as described above, the registration may be confirmed by a user operating the client device 330.
[0156] As illustrated by operation 380, the server 340 determines image pairs based on the server 340 registering the OCT B-scan images. As noted above, the server 340 may determine a correspondence between OCT B-scan images captured at different points in time, based on the registration of corresponding en-face images of a patient’s eye at a first point in time (e.g., with respect to a first OCT volume scan) and en-face images of the patient’s eye at a second point in time (e.g., with respect to a second OCT volume scan). The server 340 may then determine one or more correspondence between pairs of one or more OCT B-scans included in a first OCT volume scan and a second OCT volume scan and, based on the correspondence, the server 340 may determine the image pairs.
[0157] As illustrated by operation 382, the server 340 trains a model 344 based on the OCT B-scan image pairs. In some implementations, the model 344 is the same as, or similar to, the model 342. During training, the model 344 may be trained as follows. For a number of training steps (e.g., 40,000 training steps), provide an OCT B-scan from a first OCT volume image captured at a first point in time to a first layer of a GAN. The GAN may include a generator network configured to receive the OCT B-scan image (e.g., a 512x512 image), analyze the image, and provide outputs to successive layers of the GAN as described herein. In some embodiments, the generator network of the GAN may be configured to output an image (e.g., a 512x512 image) representing a predicted OCT B-scan image of the patient’s eye in the future (e.g., two years from the date associated with the first OCT volume image). In some embodiments, the GAN may include a discriminator network. The discriminator network of the GAN may be configured to receive the image output by the generator network and the second OCT B-scan (e.g., a target image representing the actual state of the patient’s eye two years later). The discriminator network may be configured to output a determination about whether the image that was output by the generator network is real or fake. The resulting model 344 may then be used as described with respect to FIGS. 4A-4D to generate predicted images.
[0158] FIGS. 4A-4D illustrate a non-limiting example of an implementation 400 of techniques for generating predicted images of patient retinas at inference time in accordance with one or more embodiments. In some implementations, one or more of the computing devices and / or models described may be the same as, or similar to, one or more of the computing devices and / or models of FIG. 1 or FIGS. 3A-3H.
[0159] Referring to FIG. 4A, as illustrated by an operation 460, an OCT device 410 generates an image. The OCT device 410 may be the same as, or similar to, OCT devices 110 of FIG. 1. When generating the image, the OCT device 410 may generate data associated with a standard OCT image. Additionally, the OCT device 410 may generate a plurality of images, where each image represents at least a portion of a retina of a patient.
[0160] As illustrated by an operation 462, the OCT device 410 transmits the data associated with the image to a server 440. The server 440 may be the same as, or similar to, the server 140 of FIG. 1. In some implementations, the server 440 stores the data associated with the image in a database (not explicitly illustrated). In examples, the server 440 may receive multipleimages associated with the patient involved in the image over time, and the server may store the multiple images in the database. In this example, in response to receiving requests from a client device 430 (shown in FIG. 4D) for data associated with one or more images of the retina of the patient, the server 440 may retrieve the data associated with the one or more images and transmit the data to the client device 430. The client device 430 may then cause a display associated with the client device (not explicitly illustrated) to display the one or more images based on (e.g., in response to) receiving the data associated with the one or more images.
[0161] As illustrated by an operation 464, the server 440 pre-processes the image. For example, the server 440 may pre-process the image by cropping the image (e.g., going from 1024x512 to 900x512). Additionally, or alternatively, the server 440 may resize the image (e.g., going from 1024x512 or 900x512 to 512x512). In some examples, the server 440 may pre-process the image by matching a histogram of the image to an average histogram. In this example, the average histogram may be associated with (e g., correspond to) an average histogram of images used to train a model involved in generating predicted images.
[0162] Referring to FIG. 4B, as illustrated by an operation 466, the server 440 provides the image (the pre-processed image) to a model (e.g., a model that is the same as, or similar to, the model 344 of FIG. 3H). For example, the server 440 may provide the image to a model trained in accordance with the techniques described by the present disclosure.
[0163] As illustrated by an operation 468, the server 440 generates data associated with a predicted image. In this example, the server 440 generates the data associated with the predicted image based on the server 440 providing data associated with the image to the model. In some examples, the model may be a generative model included in a generative adversarial network (GAN) that is trained to generate predicted images as described with respect to FIG. 2.
[0164] Referring to FIG. 4C, as illustrated by an operation 470, the server 440 then causes the model to generate an output. The output is associated with a predicted image representing the retina of the patient at a predetermined period of time in the future. In some examples, the predetermined period of time is two years from the point in time at which the image is generated by the OCT device 410.
[0165] As illustrated by an operation 472, the server 440 generates data associated with the predicted image. The data associated with the predicted image may be configured to cause the client device 430 to display the predicted image.
[0166] Referring to FIG. 4D, as illustrated by an operation 474, the server 440 provides (e g., transmits) the data associated with the predicted image to the client device 430. As illustrated by an operation 476, the client device 430 may display the predicted image. In some implementations, the client device 430 generates the predicted image based on receiving the data associated with the predicted image.
[0167] FIG. 5 illustrates a non-limiting example of an image pair 500 with an image 510 and a segmented image 520, in accordance with one or more embodiments described herein. More specifically, the image pair 500 includes a first image 510 and the segmented image 520 corresponding to the first image 510. The first image 510 represents a B-scan generated by an OCT device (e.g., an OCT device that is the same as, or similar to, OCT devices 110 of FIG. 1 and / or described with respect to FIG. 2 and FIGS. 3A-3D). The B-scan illustrates different layers of a retina of a patient.
[0168] With continued reference to FIG. 5, the segmented image 520 represents multiple the different layers of the retina of the patient divided by boundary lines (520a, 520b, 520c, 520d, 520e). The boundary lines 520a-520e are smoothed to enable continuity of each boundary line across the image. As illustrated, the boundary lines 520a-520e are overlaid onto the first image 510 to generate the segmented image 520. It will be understood that the boundary lines 520a-520e may also be represented as metadata. In some implementations, a server (e.g., a server that is the same as, or similar to, the server 140 of FIG. 1) may generate data associated with the first image 510 and / or the segmented image 420 that causes a client device (e.g., a client device that is the same as, or similar to, client devices 130 of FIG. 1) to display the first image 510 and / or the segmented image 520. In some implementations, when displaying the segmented image 520, a clinician may provide input via an input device (e.g., a keyboard, a mouse, and / or the like) that causes a cursor to be placed in proximity to (or over) the boundary lines 520a-520e. In this example, the data associated with the first image 510 and / or the segmented image 520 that causes a client device to display the first image 510 and / or the segmented image 520 may further causethe display to illustrate the corresponding boundary line 520a-520e to be highlighted (e.g., illustrated in a different color, illustrated as a thicker line, and / or the like).
[0169] FIG. 6 is a flow chart illustrating operations of a method 600 for detecting instances of disease signatures using a machine-learning architecture, in accordance with one or more embodiments. In some implementations, one or more of the functions described with respect to method 600 may be performed (e.g., completely, partially, and / or the like) by a server that is the same as, or similar to, server 140. In some implementations, one or more of the functions described with respect to method 600 may be performed (e.g., completely, partially, and / or the like) by another device or group of devices separate from and / or including the server, such as by one or more OCT devices that are the same as, or similar to, OCT devices 110 of FIG. 1, one or more client devices that are the same as, or similar to, the client devices 130 of FIG. 1, and / or a database that is the same as, or similar to, the database 150 of FIG. 1.
[0170] At operation 610, the method includes obtaining, by a server, volumetric data representing three-dimensional imagery. For example, the server can obtain volumetric data representing a three-dimensional image of an eyeball of a patient. The server can obtain the volumetric data based on the generation of one or more OCT B-scans generated an OCT device (e.g., that is the same as, or similar to, the OCT device described with respect to FIG. 2). In some embodiments, the OCT B-scans can be generated along a common axis or set of axis (e.g., as a set of consecutive planes extending across the eyeball of the patient).
[0171] At operation 620 the method includes identifying, by the server, one or more layer boundaries. For example, the server can identify one or more layer boundaries of the eyeball of the patient (e.g., boundary layers of portions of a retina of a patient). In this example, the computer can identify the one or more boundary layers based on (e.g., within) the plurality of OCT B-scans (also referred to as OCT B-scan cross-sectional images) of the volumetric data.
[0172] In some embodiments, the server can identify the one or more boundary layers based on (e.g., using) a model as described herein. For example, the server can identify the one or more boundary layers of a choroid layer of the retina of the eyeball based on the server providing the one or more OCT B-scans to a model configured to annotate the pixels corresponding to the RPE layer. The model may be the same as, or similar to, the model described with respect tooperation 240 of FIG. 2. For example, in the case of a GAN, the server may provide images (e.g., non-segmented OCT B-scans) to the generator network of the GAN. The generator network of the GAN can then be configured to output the image and / or annotations to be applied to the image (e.g., to the input OCT B-scans) to indicate the boundaries of the choroid layer of eyeball represented by the input image.
[0173] In some embodiments, the model used to identify the one or more boundary layers of the RPE layer of the eyeball can include a generator network of a GAN that was trained in coordination with a discriminator network of the GAN. For example, the server can provide the OCT B-scans to the generator of the GAN and corresponding ground truth image (e.g., segmented images including annotated pixels corresponding to the boundary layers of the RPE layer or vascular layer of the eyeball of the patient) to the discriminator network of the GAN. The server can then cause the model to be updated (e.g., by updating one or more weights associated with the generator network or the discriminator network) based on the output of the discriminator network and repeat this process until the GAN converges (e.g., until the generator generates images that the discriminator is unable to identify as being generated without satisfying a predetermined error threshold). The server may perform this process prior to training the model on image pairs of images captured at different points of time as described herein.
[0174] At operation 630, the server extracts or generates one or more en-face images using the B-scan images of the OCT volumetric data. The server acquires a series of B-scans from the OCT device, each B-scan represents a cross-sectional view of the retina at different lateral positions, and when combined, form the OCT volumetric dataset. The B-scans are stacked together to create a 3D representation of the retina. The volume is organized as a grid of depth-resolved images along both the lateral (X-Y) and axial (Z) dimensions. In the OCT volumetric data, each pixel or portion contains attribute data, such as information about reflectivity or pixel intensities at different layers or depths (Z-axis) of the retina corresponding to the boundaries within the volumetric data. As mentioned in FIG. 2 (at operation 210), the server may align the multiple B- scans of an OCT volume scan relative to an en-face image, which is a B-scan that is transverse to the multiple B-scans. The server may segment each layer of each image of an OCT volume scan and extract an en-face image from the OCT volume image.
[0175] Using the boundary layers (identified in prior step 620), the server executes a segmentation model trained to identify the boundary layer corresponding to a particular boundary corresponding to a type of disease indicator. The server selects a specific depth plane (a Z-level) within the volumetric data according to the particular boundary layer of the particular type of layer, where the depth corresponds to the particular retinal layer (e.g., RPE).
[0176] As an example, in some implementations (as in FIG. 7A), the machine-learning architecture is trained to detect instances of drusen as disease indicators. In such implementations, the segmentation model is trained to identify the Retinal Pigment Epithelium (RPE) layer boundary correspond to the RPE layer. As an example, in some implementations, the machine-learning architecture is trained to detect instances of geographic atrophy (GA) as disease indicators (as in FIG. 7B). In such implementations, the segmentation model is trained to identify the vascular layer boundary correspond to a vascular layer.
[0177] In some embodiments, the server can train and / or update the model used to segment the pixels corresponding to the choroid layer. For example, the server can provide one or more OCT B-scans to the generator network of the GAN to cause the generator network to generate an output. The output of the generator network can include annotations of the OCT B-scan indicative of the boundary layers of the choroid layer. The server can then provide the output of the generator network and a ground truth image to the discriminator network to cause the discriminator network to generate an output indicating whether the output of the generator network is real or not real (e.g., generated by the generator network). The server can then update the weights of the generator network and / or the discriminator network until the generator network is able to cause the discriminator network to misclassify the outputs of the generator network to satisfy a quality threshold, indicating the generator network has converged.
[0178] To generate the en-face image, the server generates an average of various attribute values (e.g., reflectivity values) of the B-scans, at the selected depth across the lateral dimensions (X-Y axes) of the B-scans. The extracted en-face image includes two-dimensional imagery from a top-down perspective at the particular retinal layer of the three-dimensional volumetric data.
[0179] At operation 640, the server determines one or more features of the en-face image that are indicative of a disease indicator. The server may execute a segmentation model using theen-face image, where the segmentation model is trained to receive the en-face image as an input, identify the features of the disease indicators, and generate an output comprising a segmented en- face image, which may include an updated en-face image based on the identified and segmented features of the disease indicator. The segmentation model is trained to detect certain features of these structures corresponding to dimensions or contours of the structures. Where the segmentation model determines these features satisfy certain segmentation thresholds or signatures, the segmentation model may identify or segment these structures as one or more disease indicators. The segmentation model then generates or otherwise outputs the segmented en-face image.
[0180] As an example, where the machine-learning architecture is trained to detect instances of drusen as disease indicators in the en-face image (as in FIG. 8A), the server extracts an en-face image using averaged values or features at the RPE layer, such that the en-face image displays structures in the RPE layer. In this example, the structures include drusen deposits that occur at pixels having reflectivity values that satisfy a segmentation threshold.
[0181] As another example, where the machine-learning architecture is trained to detect instances of GA as disease indicators in the en-face image (as in FIG. 8B), the server extracts an en-face image using averaged values or features at the vascular layer, such that the en-face image displays structures in the vascular layer. In this example, the structures include distortions or shapes in the vascular layer that occurs at pixels having reflectivity values that satisfy a segmentation threshold.
[0182] At operation 650, the server generates the updated en-face image as the segmented image, using the segmentation model. The server may output the updated en-face image for display at a user interface of a user, such as a clinician or patient. The server may execute a disease prediction machine-learning model for detecting a disease likelihood or disease indicators, using the updated en-face image. The server may generate en-face image to include highlighting or otherwise indicate one or more aspects represented in the en-face image that is indicative of a particular disease for which the segmentation model is trained to identify.
[0183] FIGS. 7A-7B show depict dataflow of processes 700a-700b for implementing a segmentation model 742a-742b (generally referred to as a segmentation model 742) of a machinelearning architecture, where the segmentation model 742 includes a machine-learning modelprogrammed and trained to ingest one or more B-scan cross-sectional input image 710a-710b (generally referred to an input image 710) and output a segmented image 720a-720b (generally referred to as a segmented image 720), in accordance with one or more embodiments described herein. In a training phase, a server (or other computer) executing the segmentation model 742 may reference ground truth, segmented, target images 725a-725b (generally referred to as target images 725).
[0184] At training time, the server trains the segmentation model 742 based on the input images 710 and target segmented images (target images 725), as similarly described in FIG. 2 and FIGS. 3C-3D. The input image 710 represents a B-scan generated by an OCT device (e.g., an OCT device 110 of FIG. 1). The B-scan of the input image 710 illustrates different layers of a retina of a patient. The segmented image 720 represents multiple the different retina layers 750 of the retina as detected and identified by the model 742 according to boundary lines (as similarly described in the image pair 500 of FIG. 5). In some implementations, the segmentation model 742 is a GAN as described herein (e.g., a Pix2Pix GAN and / or the like). The server obtains one or more training input images 710. In some cases, the server receives the training input images 710 (e g., from a database, from an administrative client device). In some cases, the server receives the training input images 710 including training B-scans of training volumetric data. The server may further receive certain types of training data, including the target images 725 and training metadata indicating structures or features in the image data of the target images 725 associated with or otherwise indicating the structures to be identified and segmented by the segmentation model 742.
[0185] For a number of training steps (e.g., 40,000 training steps), the server executes certain layers of the segmentation model 742 for ingesting the input image 710. As an example, the GAN may include a generator network configured to receive a histogram mapped segmented image 720 (e.g., a 512x512 image), analyze the segmented image 720, and provide outputs to successive layers of the GAN as described herein. In some embodiments, the generator network of the GAN may be configured to output a segmented image 720 (e.g., a 512x512 image) having one or more boundary lines. In some embodiments, the GAN may include a discriminator network. The discriminator network of the GAN may be configured to receive the segmented image 720, as output by the generator network, and a boundary image (e.g., a target image 725 representing one or more boundary lines of structures). The discriminator network may be configured to output adetermination about whether the segmented image 720 that was output by the generator network is real or fake according to the corresponding target image 725. During training, the server may calculate the cross entropy (loss) (e.g., level of error of the segmentation model 742) between the predicted boundaries of structures of the segmented image 720 and the boundaries of structures in the target image 725. Based on the computed loss, the loss function or other function of the segmentation model 742 updates one or more weights or parameters of the model segmentation model 742 to reduce the loss. The server may repeat the training process until the segmentation model 742 converges, where the loss satisfies a training threshold.
[0186] In some implementations, a server (e.g., a server that is the same as, or similar to, the server 140 of FIG. 1) may generate data associated with the input image 710 and / or the segmented image 720 that causes a client device (e.g., a client device that is the same as, or similar to, client devices 130 of FIG. 1) to display the input image 710 and / or the segmented image 720. In some implementations, when displaying the segmented image 720, a clinician may provide input via an input device (e.g., a keyboard, a mouse, and / or the like) that causes a cursor to be placed in proximity to (or over) the boundary lines. In some cases, the data associated with the input image 710 and / or the segmented image 720 that causes a client device to display the input image 710 and / or the segmented image 720 may further cause the display to illustrate the corresponding boundary line to be highlighted (e.g., illustrated in a different color, illustrated as a thicker line, and / or the like).
[0187] FIG. 7A depicts the dataflow of a process 700a for implementing a machinelearning model 742a trained to ingest an input image 710a and output a segmented image 720a according to instances of drusen structures 740. The server trains the segmentation model 742b using ground truth training data to identify the layer boundaries and identify the RPE layer 760 within the different retina layers 750. The drusen structures 740 include deposits in the eyeball, where the pixels of the input image 710a at the RPE layer 760 having reflectivity values and contours that, for example, satisfy segmentation signatures or thresholds for drusen structures 740.
[0188] FIG. 7B depicts the dataflow of a process 700b for implementing a machinelearning model 742b trained to ingest an input image 710b and output a segmented image 720b according to instances of GA structures 780. The server trains the segmentation model 742b using ground truth training data to identify the layer boundaries and identify the vascular layer 770 withinthe different retina layers 750. The GA structures 780 include distortions in the eyeball at the vascular layer 770 of the eyeball, where the pixels of the input image 710b at the vascular layer 770 have reflectivity values and contours that, for example, satisfy segmentation signatures or thresholds for GA structures 780.
[0189] FIGS. 8A-8B depict dataflow of processes 800a-800b for training or implementing a segmentation model 842a-742b (generally referred to as a segmentation model 842) of a machine-learning architecture, where the segmentation model 842 is programmed and trained to ingest one or more en-face input images 810a-810b (generally referred to an en-face input image 810) and output an en-face segmented image 820a-820b (generally referred to as an en-face segmented image 820), in accordance with one or more embodiments described herein. In a training phase, a server (or other computer) executing the segmentation model 842 may reference ground truth, segmented, target images 825a-825b (generally referred to as target images 825).
[0190] At training time, the server trains the segmentation model 842 based on the en-face input images 810 and target segmented images (target images 825), as similarly described in FIG. 2 and FIGS. 3C-3D. In some implementations, the segmentation model 842 is a GAN as described herein (e.g., a Pix2Pix GAN and / or the like). The server obtains one or more training en-face input images 810. In some cases, the server receives the training en-face input images 810 (e g., from a database, from an administrative client device). In some cases, the server extracts the training en-face input images 810 from training B-scans of training volumetric data. The server may further receive certain types of training data, including the target images 825 and training metadata indicating structures or features in the image data of the target images 825 associated with or otherwise indicating the structures to be identified and segmented by the segmentation model 842.
[0191] For a number of training steps (e.g., 40,000 training steps), the server executes certain layers of the segmentation model 842 for ingesting the en-face input image 810. As an example, the GAN may include a generator network configured to receive a histogram mapped en-face segmented image 820 (e g., a 512x512 image), analyze the en-face segmented image 820, and provide outputs to successive layers of the GAN as described herein. In some embodiments, the generator network of the GAN may be configured to output an en-face segmented image 820 (e.g., a 512x512 image) having one or more boundary lines. In some embodiments, the GAN mayinclude a discriminator network. The discriminator network of the GAN may be configured to receive the en-face segmented image 820, as output by the generator network, and a boundary image (e.g., a target image 825 representing one or more boundary lines of structures). The discriminator network may be configured to output a determination about whether the en-face segmented image 820 that was output by the generator network is real or fake according to the corresponding target image 825. During training, the server may calculate the cross entropy (loss) (e.g., level of error of the segmentation model 842) between the predicted boundaries of structures of the en-face segmented image 820 and the boundaries of structures in the target image 825. Based on the computed loss, the loss function or other function of the segmentation model 842 updates one or more weights or parameters of the model segmentation model 842 to reduce the loss. The server may repeat the training process until the segmentation model 842 converges, where the loss satisfies a training threshold.
[0192] In some implementations, a server (e.g., a server that is the same as, or similar to, the server 140 of FIG. 1) may generate data associated with the en-face input image 810 and / or the en-face segmented image 820 that causes a client device (e.g., a client device that is the same as, or similar to, client devices 130 of FIG. 1) to display the en-face input image 810 and / or the en-face segmented image 820. In some implementations, when displaying the en-face segmented image 820, a clinician may provide input via an input device (e.g., a keyboard, a mouse, and / or the like) that causes a cursor to be placed in proximity to (or over) the boundary lines. In some cases, the data associated with the en-face input image 810 and / or the en-face segmented image 820 that causes a client device to display the en-face input image 810 and / or the en-face segmented image 820 may further cause the display to illustrate the corresponding boundary line to be highlighted (e.g., illustrated in a different color, illustrated as a thicker line, and / or the like).
[0193] FIG. 8A depicts the dataflow of a process 800a for implementing a segmentation model 842a trained to ingest an en-face input image 810a and output an en-face segmented image 820a according to instances of drusen structures. The server trains the segmentation model 842b using ground truth training data (en-face target images 825a) to identify the layer boundaries and identify the RPE layer within the different retina layers. The drusen structures include deposits in the eyeball, where the pixels of the en-face input image 810a at the RPE layer having reflectivityvalues and contours that, for example, satisfy segmentation signatures or thresholds for drusen structures.
[0194] FIG. 8B depicts the dataflow of a process 800b for implementing a segmentation model 842b trained to ingest en-face input image 810b and output an en-face segmented image 820b according to instances of GA structures. The server trains the segmentation model 842b using ground truth training data (en-face target images 825b) to identify the layer boundaries and identify the vascular layer within the different retina layers. The GA structures include distortions in the eyeball at the vascular layer of the eyeball, where the pixels of the en-face input image 810b at the vascular layer have reflectivity values and contours that, for example, satisfy segmentation signatures or thresholds for GA structures.
[0195] FIG. 9 is a flow chart illustrating operations of a method 900 for analyzing choroid layer using a machine-learning architecture, in accordance with one or more embodiments. In some implementations, one or more of the functions described with respect to method 900 may be performed (e.g., completely, partially, and / or the like) by a server that is the same as, or similar to, server 140. In some implementations, one or more of the functions described with respect to method 200 may be performed (e.g., completely, partially, and / or the like) by another device or group of devices separate from and / or including the server, such as by one or more OCT devices that are the same as, or similar to, OCT devices 110 of FIG. 1, one or more client devices that are the same as, or similar to, the client devices 130 of FIG. 1, and / or a database that is the same as, or similar to, the database 150 of FIG. 1.
[0196] At operation 910, the server obtains volumetric data representing three-dimensional imagery, based upon image data of OCT B-scans received from an OCT device. For instance, the server can obtain the volumetric data representing a three-dimensional image of an eyeball of a patient. The server can obtain the volumetric data based on the generation of one or more OCT B- scans generated an OCT device (e.g., that is the same as, or similar to, the OCT device described with respect to FIG. 2). In some embodiments, the OCT B-scans can be generated along a common axis or set of axes (e.g., as a set of consecutive planes extending across the eyeball of the patient).
[0197] At operation 920, the server identifies one or more layer boundaries corresponding to retinal layers of the eyeball. The computer can identify the one or more boundary layers basedon (e.g., within) the plurality of OCT B-scans (also referred to as OCT B-scan cross-sectional images) of the volumetric data. In some embodiments, the server can identify the one or more boundary layers based on (e.g., using) a model as described herein. For example, the server can identify the one or more boundary layers of a choroid layer of the retina of the eyeball based on the server providing the one or more OCT B-scans to a model configured to annotate the pixels corresponding to the choroid layer. The choroid layer includes an inner surface and outer surface, such that the choroid layer includes two boundaries. The server may execute a segmentation model of a machine-learning architecture trained to identify one or both of the boundary layers of the choroid layer. An example of a model that can be used to identify the one or more boundary layers within a given OCT B-scan is described with respect to the model architecture 1300 of FIGS. 13A and 13B.
[0198] In some embodiments, the server can train and / or update the model used to segment the pixels corresponding to the choroid layer. For example, the server can provide one or more OCT B-scans to the generator network of the GAN to cause the generator network to generate an output. The output of the generator network can include annotations of the OCT B-scan indicative of the boundary layers of the choroid layer. The server can then provide the output of the generator network and a ground truth image to the discriminator network to cause the discriminator network to generate an output indicating whether the output of the generator network is real or not real (e.g., generated by the generator network). The server can then update the weights of the generator network and / or the discriminator network until the generator network is able to cause the discriminator network to misclassify the outputs of the generator network to satisfy a quality threshold, indicating the generator network has converged.
[0199] At operation 930, the server generates a three-dimensional representation of each vessel in the choroid layer of the eyeball. The server can generate the three-dimensional representation of each vessel in the choroid layer based on one or more OCT B-scans. For example, the server can identify the one or more boundary layers of each vessel of the choroid layer of the retina of the eyeball based on the server providing the one or more OCT B-scans to a model configured to annotate the pixels corresponding to the boundary (e.g., the edges) of the vessels represented by the input OCT B-scans. In some embodiments, the model can be the same as, or similar to, the segmentation models described herein. For example, the model can include a GANhaving a generator network and / or a discriminator network. The GAN can be configured to receive the OCT B-scans and use the generator network to generate annotations and / or output images that include annotations indicating the boundaries of the vessels in the choroid layer of the eyeball of the patient using the generator network. For example, the server can provide the OCT B-scan to the generator network of GAN to cause the generator network to generate an output, where the output includes annotations that indicate the edges of the vessels within the choroid layer. The server can then stack the segmented OCT B-scans and generate a three-dimensional representation of each vessel.
[0200] In some embodiments, the server can train and / or update the model used to segment the pixels corresponding to the boundaries of the vessels in the choroid layer. For example, the server can provide one or more OCT B-scans to the generator network of the GAN to cause the generator network to generate an output. The output of the generator network can include annotations of the OCT B-scan indicative of the boundary layers of the choroid layer. The server can then provide the output of the generator network and a ground truth image to the discriminator network to cause the discriminator network to generate an output indicating whether the output of the generator network is real or not real (e.g., generated by the generator network). The server can then update the weights of the generator network and / or the discriminator network until the generator network is able to cause the discriminator network to misclassify the outputs of the generator network to satisfy a quality threshold, indicating the generator network has converged.
[0201] In some embodiments, the server can generate a point cloud representing the vessels of the choroid layer based on the stacked OCT B-scans that include the annotations indicating the edges of the vessels. For example, the server can determine a centerline that extends along each vessel by fitting the centerline (e.g., a line defined by two points, a spline, and / or the like) to each vessel represented in each OCT B-scan. The server can then determine cross sections that extend substantial transverse to the one or more points of the centerline (e.g., along cross-sections associated with the centerline) and fit a circle between the boundary points (e.g., edges of each vessel where they intersect with the centerline to determine a three-dimensional surface for the portion of the vessel at that point of the centerline. The server can iteratively repeat this process for each vessel in the OCT B-scans until three-dimensional surfaces are constructed for each vessel of the plurality of OCT B-scans. As will be understood, by iteratively repeating this process offitting circles between the boundary points for each centerline, the server can generate a vector image representation of each vessel.
[0202] At operation 940, the server generates a heatmap based on the three-dimensional representation of each vessel, the heatmap representing variation in diameter of the vessels of the eyeball along centerlines defined by the vessels. For example, the server can determine an intensity (e.g., when generating a greyscale heatmap) or a color along a gradient, where the intensity or color correspond to the diameter of the vessel at each point along the vessels. Once intensities or colors are assigned to the portions of each vessel, the server can generate a GUI based on the intensity or color assigned to the portions of each vessel. The server can then generate GUI data associated with the GUI and provide the GUI data (e.g., as part of an output file) to a client device to cause the client device to display the GUI.
[0203] FIG. 10 depicts a dataflow of a portion of process (e.g., that is the same as, or similar to, the method 900 of FIG. 9) for implementing a segmentation model 1042, where the segmentation model 1042 is trained to ingest one or more OCT B-scans 1010 and output corresponding segmented images 1020, in accordance with one or more embodiments described herein. The segmentation model 1042 may be trained as similarly described herein at, without limitation, FIG. 2 and FIGS. 3C and 3D. The input image 1010 represents an OCT B-scan generated by an OCT device (e.g., an OCT device that is the same as, or similar to, the OCT device 110 of FIG. 1). The input images 1010 can illustrate different layers of a retina of a patient. The segmented images 1020 can represent multiple layer boundaries of retina layers 1050, including a choroidal inner boundary and a choroidal outer boundary that bound a choroid layer. In a training phase, a server (or other computer) executing the segmentation model 1042 may reference ground truth, segmented, target images 1025.
[0204] At training time, the server trains the segmentation model 1042 based on the input images 1010 and target segmented images (target images 1025), as similarly described in FIG. 2 and FIGS. 3C-3D. The input image 1010 represents a B-scan generated by an OCT device (e.g., an OCT device 110 of FIG. 1). The B-scan of the input image 1010 illustrates different layers of a retina of a patient. The segmented image 1020 represents multiple the different retina layers 1050 corresponding to the choroid layer of the retina, as detected and identified by the model 1042, according to boundary lines (as similarly described in the image pair 500 of FIG. 5). In someimplementations, the segmentation model 1042 is a GAN as described herein (e g., aPix2Pix GAN and / or the like). The server obtains one or more training input images 1010. In some cases, the server receives the training input images 1010 (e.g., from a database, from an administrative client device). In some cases, the server receives the training input images 1010 including training B- scans of training volumetric data. The server may further receive certain types of training data, including the target images 1025 and training metadata indicating structures or features in the image data of the target images 1025 associated with or otherwise indicating the structures to be identified and segmented by the segmentation model 1042.
[0205] For a number of training steps (e.g., 40,000 training steps), the server executes certain layers of the segmentation model 1042 for ingesting the input image 1010. As an example, the GAN may include a generator network configured to receive a histogram mapped segmented image 1020 (e.g., a 512x512 image), analyze the segmented image 1020, and provide outputs to successive layers of the GAN as described herein. In some embodiments, the generator network of the GAN may be configured to output a segmented image 1020 (e.g., a 512x512 image) having one or more boundary lines. In some embodiments, the GAN may include a discriminator network. The discriminator network of the GAN may be configured to receive the segmented image 1020, as output by the generator network, and a boundary image (e.g., a target image 1025 representing one or more boundary lines of structures). The discriminator network may be configured to output a determination about whether the segmented image 1020 that was output by the generator network is real or fake according to the corresponding target image 1025. During training, the server may calculate the cross entropy (loss) (e.g., level of error of the segmentation model 1042) between the predicted boundaries of structures of the segmented image 1020 and the boundaries of structures in the target image 1025. Based on the computed loss, the loss function or other function of the segmentation model 1042 updates one or more weights or parameters of the model segmentation model 1042 to reduce the loss. The server may repeat the training process until the segmentation model 1042 converges, where the loss satisfies a training threshold.
[0206] In some implementations, a server (e.g., a server that is the same as, or similar to, the server 140 of FIG. 1) may generate data associated with the input image 1010 and / or the segmented image 1020 that causes a client device (e.g., a client device that is the same as, or similar to, client devices 130 of FIG. 1) to display the input image 1010 and / or the segmented image 1020.In some implementations, when displaying the segmented image 1020, a clinician may provide input via an input device (e.g., a keyboard, a mouse, and / or the like) that causes a cursor to be placed in proximity to (or over) the boundary lines. In some cases, the data associated with the input image 1010 and / or the segmented image 1020 that causes a client device to display the input image 1010 and / or the segmented image 1020 may further cause the display to illustrate the corresponding boundary line to be highlighted (e.g., illustrated in a different color, illustrated as a thicker line, and / or the like).
[0207] FIG. 11 depicts a dataflow of a portion of process 1100 (e.g., that is the same as, or similar to, the method 900 of FIG. 9) for implementing a segmentation model 1142 programmed and trained to ingest one or more OCT B-scans 1110 and output corresponding segmented images 1120 for vessel segments 1140 in a choroid layer 1150, in accordance with one or more embodiments described herein. The segmentation model 1142 may be trained as similarly described herein at, without limitation, FIG. 2 and FIGS. 3C-3D. The input image 1110 represents an OCT B-scan generated by an OCT device (e.g., an OCT device that is the same as, or similar to, the OCT device 110 of FIG. 1). The input images 1110 can illustrate different layers and structures of a retina of a patient, including layer boundaries of the choroid layer 1150 and vessels 1140 of the choroid layer 1150.
[0208] At training time, the server trains the segmentation model 1142 based on the input images 1010 and target segmented images (target images 1125), as similarly described in FIG. 2 and FIGS. 3C-3D. The input image 1110 represents a B-scan generated by an OCT device (e.g., an OCT device 110 of FIG. 1). The B-scan of the input image 1110 illustrates different layers of a retina of a patient. The segmented image 1120 represents multiple the different retina layers corresponding to the choroid layer 1150 of the retina and the vessels 1140, as detected and identified by the model 1142, according to boundary lines (as similarly described in the image pair 500 of FIG. 5) of the layers and structures (e.g., boundary lines of the vessels 1140). In some implementations, the segmentation model 1142 is a GAN as described herein (e.g., a Pix2Pix GAN and / or the like). The server obtains one or more training input images 1110. In some cases, the server receives the training input images 1110 (e.g., from a database, from an administrative client device). In some cases, the server receives the training input images 1110 including training B- scans of training volumetric data. The server may further receive certain types of training data,including the target images 1125 and training metadata indicating structures or features in the image data of the target images 1125 associated with or otherwise indicating the structures to be identified and segmented by the segmentation model 1042. For instance, the target images 1125 may indicate layer boundaries of the choroid layer 1150 and the structure boundaries of the vessels 1140 within the choroid layer 1150.
[0209] For a number of training steps (e.g., 40,000 training steps), the server executes certain layers of the segmentation model 1142 for ingesting the input image 1110. As an example, the GAN may include a generator network configured to receive a histogram mapped segmented image 1120 (e.g., a 512x 12 image), analyze the segmented image 1120, and provide outputs to successive layers of the GAN as described herein. In some embodiments, the generator network of the GAN may be configured to output a segmented image 1120 (e.g., a 512x512 image) having one or more boundary lines of the layers or the structures. In some embodiments, the GAN may include a discriminator network. The discriminator network of the GAN may be configured to receive the segmented image 1120, as output by the generator network, and a boundary image (e.g., a target image 1125 representing one or more boundary lines of structures). The discriminator network may be configured to output a determination about whether the segmented image 1120 that was output by the generator network is real or fake according to the corresponding target image 1125. During training, the server may calculate the cross entropy (loss) (e.g., level of error of the segmentation model 1142) between the predicted boundaries of the choroid layer 1150 and the vessels 1140 structures of the segmented image 1120 and the corresponding boundaries of the choroid layer 1150 and the vessels 1140 structures in the target image 1125. Based on the computed loss, the loss function or other function of the segmentation model 1142 updates one or more weights or parameters of the model segmentation model 1142 to reduce the loss. The server may repeat the training process until the segmentation model 1142 converges, where the loss satisfies a training threshold.
[0210] In some implementations, a server (e.g., a server that is the same as, or similar to, the server 140 of FIG. 1) may generate data associated with the input image 1110 and / or the segmented image 1120 that causes a client device (e.g., a client device that is the same as, or similar to, client devices 130 of FIG. 1) to display the input image 1110 and / or the segmented image 1120. In some implementations, when displaying the segmented image 1120, a clinician may provideinput via an input device (e.g., a keyboard, a mouse, and / or the like) that causes a cursor to be placed in proximity to (or over) the boundary lines. In some cases, the data associated with the input image 1110 and / or the segmented image 1120 that causes a client device to display the input image 1110 and / or the segmented image 1120 may further cause the display to illustrate the corresponding boundary lines of the choroid layer 1150 and / or the vessels 1140 to be highlighted (e.g., illustrated in a different color, illustrated as a thicker line, and / or the like).
[0211] FIG. 12 depicts operations and dataflow of a method 1200 of analyzing a choroid layer of a retina, according to some embodiments. In some implementations, one or more of the functions described with respect to method 1200 may be performed (e.g., completely, partially, and / or the like) by a server that is the same as, or similar to, server 140. In some implementations, one or more of the functions described with respect to method 1200 may be performed (e.g., completely, partially, and / or the like) by another device or group of devices separate from and / or including the server, such as by one or more OCT devices that are the same as, or similar to, OCT devices 110 of FIG. 1, one or more client devices that are the same as, or similar to, the client devices 130 of FIG. 1, and / or database that is the same as, or similar to, the database 150 of FIG. 1.
[0212] The server executes various operations, such as programming operations of a machine-learning model (e.g., segmentation model 1042, 1142) for identifying and segmenting boundaries of choroidal layers and vessel structures in the OCT B-scan images of the OCT volumetric data. The server then generates segmented images (e.g., segmented image 1020, 1120) containing visual or metadata indicators of the choroidal layers and the vessels. After the server identifies and generates the segments of the choroidal layers from the OCT B-scans of the volumetric data, the server identifies and segments the vessels in the cross sections depicted in the OCT B-scans.
[0213] After the server identifies and segments the vessels (in the segments of the choroidal layers), the server generates, creates, or otherwise obtains a choroidal vasculature model 1205. The choroidal vasculature model 1205 includes a three-dimensional image or data structure, which the server generated by stacking the choroidal segments and vessel segments from the cross- sectional OCT B-scans of the OCT volumetric data.
[0214] At operation 1210, the server selects vessel components as a portion of the choroidal vasculature model 1205 and detects centerlines of the selected vessel component, to generate the vessel centerline image 1215. Using the segment data of the choroidal vasculature model 1205, the server identifies a centerline and centerline point of the selected portions of the choroidal vasculature model 1205 containing the vessels.
[0215] In some embodiments, the server executes operations of 3D tensor voting. The 3D tensor voting includes computer vision or image processing operations for identifying and extracting structures (e.g., edges, surfaces, or curves) from image data. The 3D tensor voting may operate in the 2D and is applied in the 3D data to detect the structure attributes, such as surfaces or curves, within the 3D point cloud or OCT volumetric data. The server executes the 3D tensor voting to analyze a 3D point cloud of the vessel structure (of the portion of the choroidal vasculature model 1205) to determine a direction and orientation of a vessel centerline. The server may further execute the operations of the 3D tensor voting to determine centerline points of the vessels within the portion of the 3D model of the choroidal vasculature model 1205.
[0216] At operation 1220, the server determines an orthogonal plane at a centerline point of the vessel to generate a planar image 1225. At each centerline point, the server fits the orthogonal plane to the surrounding points of the vessel. In some implementations, the server references pixels or points of a 3D point cloud of the choroidal vasculature model 1205, at the portion containing the selected vessel. The server references of the centerline image 1215 for the centerline and the centerline point to fit the orthogonal plane at the centerline point. As shown in the planar image 1225, the orthogonal plane overlaps or intersects with the corresponding 3D point cloud around the centerline and centerline point.
[0217] At operation 1230, the server determines an intersection between the plane and a region of interest around the centerline point to generate a structure estimate image 1235. For example, the server can determine the intersection between the plane and the region of interest around the centerline point, where the region of interest includes an area between the boundary edges of the vessel.
[0218] At operation 1240, the server generates a circle fitted around a diameter of the vessel at the centerline point to generate a diameter image 1245. The server references theintersection of the orthogonal plane with the vessel centerline and centerline points to fit a circle, which the server uses to determine a vessel diameter measurement at location of the intersection between the centerline point and the orthogonal plane. The server can then generate a three- dimensional representation of the vessel.
[0219] At operation 1250, the server generates a heatmap image 1255 representing the choroidal vessel radius variation along the centerline of the choroidal vasculature model 1205. The server may repeat operations 1210-1240 at each centerline point to generate the vessel diameters at multiple locations (as in operation 1240). The server may compile the vessel diameter measurements to generate the heatmap according to variations of the diameters as computed amongst the vessels. In some cases, a color variations or other visual indicators may represent variation in vessel diameter across the choroid layer.
[0220] The server may output the heatmap and various metrics, including the diameters of the vessels. For example, the server can generate a GUI as described herein of the vessels to visually indicate the shape of the vessels. In another example, the server may output a report indicating one or more metrics that are based on the shape of the vessels. For example, the server may output a report quantifying a number of vessels that have shapes indicative of one or more diseases and / or the like. The metrics may correspond to biomarker quantifications. Non-limiting examples may include choroid layer thickness maps, choroid layer vascularity index (CVI) maps, choroid layer inner / outer surface contour maps, choroid layer area / volume measurements, choroid layer vessel volume measurements, choroid layer vessel diameter measurements, and choroid layer vessel diameter heatmaps. In some cases, generating the metrics of the biomarker quantification outputs may include generating area / volume measurements of various retinal lesions including subretinal fluid (SRF), pigment epithelium detachment (PED), and intra-retinal fluid using OCT scans. In some cases, generating the metrics for the biomarker quantification outputs may include generating measurements of retinal lesions including hard exudates using color fundus (CF) images or photographs. Additionally or alternatively, the server may generate the metrics of the biomarker quantification outputs by executing software programming of a multimodal manual measurement tool for measuring thickness, area, and volume of lesions, as well as various layers of the posterior segment of the eye, among other types of outputs. The server may perform multimodal image registration including, for example, registration across various 2D retinal imagemodalities; (ii) registration of longitudinal 2D retinal images obtained from the same modality; (iii) volumetric registration of longitudinal volumes; (iv) registration of 2D image modalities such as CF, fundus autofluorescence (FAF) with 3D OCT volumes; and (v) registration of 2D and 3D retinal images with Humphrey visual field data.
[0221] FIG. 13A depicts a model architecture 1300 that can be implemented as part of a segmentation model (e.g., that is the same as, or similar to, the segmentation models described herein including, without limitation, segmentation model 1142) programmed and trained to ingest one or more OCT B-scans (e.g., that are the same as, or similar to, the OCT B-scans described herein including, without limitation, the OCT B-scans 1110) and output corresponding segmented images, in accordance with one or more embodiments described herein. In some embodiments, the segmentation model implementing the model architecture 1300 can be implemented by a server (e.g., that is the same as, or similar to, the servers described herein including, without limitation, the server 140 of FIG. 1).
[0222] In some embodiments, the model architecture 1300 can be configured to receive as input an input image. For example, the model architecture 1300 can be configured to receive as input an input image representing an OCT B-scan generated by an OCT device (e.g., an OCT device that is the same as, or similar to, the OCT device 110 of FIG. 1). The input images can illustrate different layers and structures of a retina of a patient, including layer boundaries of the choroid layer and vessels of the choroid layer.
[0223] In some embodiments, the model architecture 1300 can be implemented as a machine learning model such as a U-net (described above) configured to implement semantic segmentation operations when identifying the various layers of the retina of the patient. For example, the model architecture 1300 can include a machine learning model having an encoderdecoder structure with a bottleneck (e.g., a residual block 1304) in between, forming a U-shaped design. The encoder-decoder structure can include a plurality of encoders 1302a-1302d that extend along a contracting path 1302 of the model architecture 1300 and are configured to extract hierarchical features by progressively downsampling the input image and / or feature maps through convolutional and max-pooling layers. In this example, each encoder 1302a-1302n can have a corresponding decoder 1306a-1306n for each stage of the model architecture 1300. One example of a stage can include encoder 1302a and corresponding decoder 1306a. In some embodiments,each decoder 1306a-1306d can be configured to reconstruct the segmentation map by upsampling and combining features from corresponding encoder layers via skip connections. This model architecture 1300 can allow for the processing of both low-level spatial details and high-level contextual information are preserved.
[0224] As described above, the inputs to model architecture 1300 can include images such as OCT B-scans as described herein. For example, the input to the model architecture 1300 can include OCT B-scans that are represented as grayscale images of cross-sectional views of the retina of a patient. The server implementing the model architecture 1300 can provide the images into a first encoder 1302a in the contracting path 1302 of the model architecture 1300. In this example, each encoder 1302a-1302d can include (e.g., implement) convolutional layers that execute operations to extract feature maps at various scales (e.g., corresponding to each stage). The output from the model architecture 1300 can include a segmented image where each pixel is assigned a label corresponding to specific retinal layers or regions of interest corresponding to the input image (e.g., the input OCT B-scan). This output can be represented as a multi-channel probability map, with each channel indicating the likelihood of a pixel belonging to a particular class corresponding to a particular layer or region of the retina.
[0225] In some embodiments, the encoder blocks 1302a-1302d in the contraction path 1302 can implement multiple convolutional layers followed by rectified linear unit (ReLU) activations and max-pooling operations for downsampling. The pool map output by the encoder blocks 1302a-1302d can be transferred via skip connections to corresponding decoders 1306a- 1306d. These pool maps can represent spatially compressed but contextually rich features that aid in reconstructing of fine details by the decoders 1306a-1306d during upsampling. The encoder blocks 1302a-1302d can also transfer the feature map generated by the encoder blocks 1302a- 1302d to the next encoder block or, in the case of encoder 1302d, to the residual block 1304. In these examples, the feature map can represent progressively abstracted features that capture higher-level patterns represented by the input image.
[0226] In some embodiments, the residual block 1304 in model architecture 1300 can include two paths: a forward path and a skip connection. In the forward path, the feature map from the last encoder 1302d in the contraction path 1302 can be processed. This can include the feature map undergoing two sequential 3x3 convolutional layers, each followed by batch normalizationand ReLU activation. In examples, the skip connection can bypass these operations and directly add the input to the output of the forward path. This can allow the residual block 1304 to learn residual mappings, simplifying the optimization process and mitigating issues such as vanishing gradients. If the input and output dimensions differ, a 1 x1 convolution can be applied to the skip connection to match dimensions before addition. In some embodiments, the output of the residual block 1304 can represent a combination of learned features from the forward path and preserved original information from the skip connection. This output can be provided to the first decoder 1306d in the expansion path 1306, where it serves as an input for upsampling operations. By retaining both high-level abstracted features and spatially detailed information, this output can allow for accurate reconstruction of segmented images during decoding.
[0227] In some embodiments, the decoders 1306a-1306d in the expansion path 1306 can execute concatenate operations to combine upsampled feature maps with skip-connected pool maps obtained by the decoders 1306a-1306n. In some examples, each decoder 1306a-1306d can then apply transposed convolutions to increase spatial resolution while reducing channel depth, followed by convolutional layers to merge contextual and localized information. The decoders 1306d-1306b can provide refined feature maps to subsequent decoders along the expansion path 1306 until reaching the output resolution. In some embodiments, the decoder 1306a (e.g., the final decoder) can output segmentation probabilities through a 1x1 convolution layer, producing masks that align with the input dimensions. The output of the decoder 1306a can represent an image (e.g., a segmented image) that corresponds to the input image where pixels representing various layers or regions are annotated using one or more channels, one or more visual indicators including predetermined colors for various layers, etc.
[0228] During training, the server can provide preprocessed OCT B-scans to the model architecture 1300 as input to cause the components of the model architecture 1300 to generate the output image as described herein. The server can then compare the output image to a ground truth image including one or more segmentation masks to calculate a difference between the output image and the ground truth image. The server can implement a loss function, such as cross-entropy, to quantify the difference to establish a degree of error between the output of the model architecture 1300 and the ground truth. In some embodiments, the server can implement backpropagation when updating (e.g., adjusting) the weights of the various components of the model architecture 1300by calculating gradients of the loss with respect to each weight. These gradients can be scaled by a learning rate and used to iteratively adjust weights, minimizing segmentation errors over successive training iterations. This process of providing input images to the model architecture 1300, comparing the input images to ground truth images, calculating corresponding losses, and updating the weights of the components of the model architecture 1300 can be iteratively repeated until a threshold difference representing convergence is achieved.
[0229] FIG. 13B depicts an architecture of an encoder 1302a. While the structure of the encoder is described with respect to the encoder 1302a, it will be understood that the architecture can be similar to or the same as that of the encoders 1302a-1302d along the contracting path of the model architecture 1300.
[0230] In some embodiments, the encoder 1302a (e.g., in the contracting path 1302 the model architecture 1300) can be designed with various configurations such as convolutional blocks or residual blocks for feature extraction and downsampling. In an example, the encoder 1302a can include two consecutive 3x3 convolutions followed by activation layers like ReLU, a max-pooling operation to reduce spatial resolution, and optional batch normalization. In some examples, one or more residual blocks (illustrated as residual blocks 1-6) can be inserted between the convolution and / or activation layers of the encoder 1302a to preserve gradient flow and mitigate the vanishing gradient issue.
[0231] In some embodiments, the encoder 1302a can receive an image at its input, such as an OCT B-scan, and output feature maps along with pooled feature maps for subsequent encoders in the contracting path 1302 or for a residual block 1304 at the end of the contraction path. In an example, the encoder 1302a can execute operations as described based on (e.g., on) the input image and its output can be passed on to the next encoder (e.g., encoder 1302b of FIG. 13A) for deeper feature extraction. In some examples, the output of the encoder 1302a can be processed using a residual block before transitioning to the expansion path. In other examples, intermediate outputs can also be kept available for skip connections to corresponding decoders as described herein.
[0232] FIG. 14A depicts operations and dataflow of a method 1400 of analyzing a choroid layer of a retina, according to some embodiments. In some implementations, one or more of the functions described with respect to method 1400 may be performed (e.g., completely, partially,and / or the like) by a server that is the same as, or similar to, server 140. In some implementations, one or more of the functions described with respect to method 1400 may be performed (e.g., completely, partially, and / or the like) by another device or group of devices separate from and / or including the server, such as by one or more OCT devices that are the same as, or similar to, OCT devices 110 of FIG. 1, one or more client devices that are the same as, or similar to, the client devices 130 of FIG. 1, and / or database that is the same as, or similar to, the database 150 of FIG. 1. In some embodiments, one or more of the functions described with respect to the method 1400 may be the same as, or similar to, those described with respect to method 1200 of FIG. 12.
[0233] In some embodiments, the server can execute operations to implement one or more of the segmentation models described herein. For example, the server can execute operations to implement a segmentation model when identifying and segmenting boundaries of choroidal layers and vessel structures of a patient's retina. The server can segment the vascular and stromal regions of the choroid. In an example, the server can process OCT B-scan images within OCT volumetric data (e.g., within a subset of the volumetric data) using a segmentation model to identify and segment the boundaries of choroidal layers and vessel structures (referred to as choroidal vasculature) of the patient's retina. In some embodiments, the server can then identify and generate the segments of the choroidal vasculature from the OCT B-scans of the volumetric data based on the segments and boundaries of the choroidal vasculature.
[0234] At operation 1410, the server can select a vessel location. For example, the server can select the vessel location within the three-dimensional image or data structure representing the choroidal vasculature. In an example, the server can select the vessel location based on input from a clinician provided to the server. In another example, the server can select the vessel location based on one or more predetermined configurations (e.g., default configurations, etc.) indicating the vessel location. At operation 1412, the server can select vessel components as a portion of the choroidal vasculature for further analysis, including centerline identification. For example, the server can select the vessel components as a portion of the choroidal vasculature and analyze the vessel components to identify a centerline. The server can then identify a centerline extending longitudinally through a vessel as shown in the planar image 1420 of FIG. 14B, as described below.
[0235] In some embodiments, the server can execute one or more operations to identify the centerline of a given vessel or vessels. For example, the server can execute a density based clustering process (e.g., Density-Based Spatial Clustering of Applications with Noise (DBSCAN), etc.) for centerline identification. In this example, the server can first processes the choroidal vasculature model (e.g., the point cloud) in slices along the y-axis to identify centerlines in vessels represented as clusters in a point cloud. This can involve dividing the point cloud into thin, horizontal slices, which allows for analysis of the vessel components in discrete segments. The server can then identify vessel centers in each slice by analyzing histogram peaks in the x-direction of each horizontal slice. This step can involve creating a histogram of the x-coordinates of the points in each slice and identifying the peaks, which correspond to the central points of the vessels in that slice. After identifying the vessel centers in each slice, the server can cluster the identified points across slices using a density-based clustering technique (e.g., DBSCAN, etc ). As a result, the server can identify densely packed points as clusters based on their spatial proximity, effectively identifying the vessel structures within the point cloud. In some examples, the server can then fit cubic splines (using a method such as splprep) to the clustered points to create smooth centerlines. This can involve interpolating the clustered points with cubic splines to generate continuous and smooth representations of the vessel centerlines.
[0236] At operation 1414, the server can determine an orthogonal plane at one or more centerline points established by the centerline of a given vessel to generate a planar image (e.g., the planar image 1420 of FIG. 14B). At the centerline point, the server can fit the orthogonal plane to the surrounding points of the vessel. In some implementations, the server can reference pixels or points of the choroidal vasculature model, at the portion containing the selected vessel. The server references of the centerline image 1420 for the centerline and the centerline point to fit the orthogonal plane at the centerline point. As shown in the planar image 1225, the orthogonal plane overlaps or intersects with the corresponding 3D point cloud around the centerline and centerline point.
[0237] At operation 1416, the server can identify an upper boundary (e.g., the retinal pigment epithelium or Bruch's membrane complex) and a lower boundary (e.g., the choroid- scleral interface) of the choroid. For example, the server can identify an upper boundary and a lower boundary indicating a thickness of the choroidal layer of the retina of the patient. In this example,the server can then determine a thickness between the upper boundary and the lower boundary along a portion of the retina of the patient associated with a given portion of the choroidal vasculature. The server can iteratively perform this process to determine the thickness of the choroidal layer across the retina of the patient to generate one or more choroidal thickness maps (e.g., that are the same as, or similar to, the choroid thickness map 1430 of FIG. 14C) or choroid vascularity index maps (e.g., that are the same as, or similar to, the choroid vascularity index map 1440 of FIG. 14D)
[0238] In some embodiments, the server can generate a heatmap image as described herein (e.g., similar to as described with respect to process 1200). For example, the server can generate a heat map image representing the choroidal vessel radius (e g., diameter), length, and branching pattern variation along the centerline of the choroidal vasculature model based on identifying the vessel points associated with the upper boundary in the lower boundary of the choroid for one or more vessels of the retina of the patient. The server can then analyze and output the heat map image to determine various metrics as described herein and generate a GUI to be displayed on a display device accessible by a clinician treating a patient.
[0239] In some embodiments, the server can cause one or more of the outputs described herein (e.g., the choroid thickness map, the choroid vascularity index map, the heatmap, etc.) to be output via a display device. For example, the server can cause one or more of the outputs described to be output via a display device of a client device that is the same as, or similar to, one or more of the client devices 130 of FIG. 1. In this example, the server can cause the one or more outputs to be output via a display device to indicate to a clinician operating the client device that portions of the retina of the patient have varying thicknesses as determined based on the execution of the operations described herein.
[0240] The foregoing method descriptions and the process flow diagrams are provided merely as illustrative examples and are not intended to require or imply that the steps of the various embodiments must be performed in the order presented. The steps in the foregoing embodiments may be performed in any order. Words such as “then,” “next,” etc., are not intended to limit the order of the steps; these words are simply used to guide the reader through the description of the methods. Although process flow diagrams may describe the operations as a sequential process, many of the operations can be performed in parallel or concurrently. In addition, the order of theoperations may be re-arranged. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, and the like. When a process corresponds to a function, the process termination may correspond to a return of the function to a calling function or a main function.
[0241] Some non-limiting embodiments of the present disclosure are described herein in connection with a threshold. As described herein, satisfying a threshold may refer to a value being greater than the threshold, more than the threshold, higher than the threshold, greater than or equal to the threshold, less than the threshold, fewer than the threshold, lower than the threshold, less than or equal to the threshold, equal to the threshold, and / or the like.
[0242] No aspect, component, element, structure, act, step, function, instruction, and / or the like used herein should be construed as critical or essential unless explicitly described as such. In addition, as used herein, the articles "a" and "an" are intended to include one or more items and may be used interchangeably with "one or more" and "at least one." Furthermore, as used herein, the term "set" is intended to include one or more items (e.g., related items, unrelated items, a combination of related and unrelated items, etc.) and may be used interchangeably with "one or more" or "at least one." Where only one item is intended, the term "one" or similar language is used. Also, as used herein, the terms "has," "have," "having," or the like are intended to be open ended terms. Further, the phrase "based on" is intended to mean "based at least partially on" unless explicitly stated otherwise.
[0243] The various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure.
[0244] Embodiments implemented in computer software may be implemented in software, firmware, middleware, microcode, hardware description languages, or any combination thereof. Acode segment or machine-executable instructions may represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. may be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, token passing, network transmission, etc.
[0245] The actual software code or specialized control hardware used to implement these systems and methods is not limiting. Thus, the operation and behavior of the systems and methods were described without reference to the specific software code being understood that software and control hardware can be designed to implement the systems and methods based on the description herein.
[0246] When implemented in software, the functions may be stored as one or more instructions or code on a non-transitory computer-readable or processor-readable storage medium. The steps of a method or algorithm disclosed herein may be embodied in a processor-executable software module, which may reside on a computer-readable or processor-readable storage medium. A non-transitory computer-readable or processor-readable media includes both computer storage media and tangible storage media that facilitate transfer of a computer program from one place to another. A non-transitory processor-readable storage media may be any available media that may be accessed by a computer. By way of example, and not limitation, such non-transitory processor-readable media may comprise RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other tangible storage medium that may be used to store desired program code in the form of instructions or data structures and that may be accessed by a computer or processor. Disk and disc, as used herein, include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc where disks usually reproduce data magnetically, while discs reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer- readable media. Additionally, the operations of a method or algorithm may reside as one or any combination or set of codes and / or instructions on a non-transitory processor-readable medium and / or computer-readable medium, which may be incorporated into a computer program product.
[0247] The preceding description of the disclosed embodiments is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the generic principles defined herein may be applied to other embodiments without departing from the spirit or scope of the disclosure. Thus, the present disclosure is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the following claims and the principles and novel features disclosed herein.
[0248] While various aspects and embodiments have been disclosed, other aspects and embodiments are contemplated. The various aspects and embodiments disclosed are for purposes of illustration and are not intended to be limiting, with the true scope and spirit being indicated by the following claims.
Claims
CLAIMSWhat is claimed is:
1. A computer-implemented method for choroid analysis of tomography (OCT) imaging, the method comprising: obtaining, by a computer, volumetric data representing three-dimensional imagery of an eyeball based on a plurality of cross-sectional images received from an OCT imaging device; executing, by the computer, a machine learning model configured to identify one or more layer boundaries corresponding to a choroid layer of the eyeball based on the plurality of cross- sectional images of the volumetric data; determining, by the computer, a subset of the volumetric data representing three- dimensional imagery of the choroid layer in response to identifying the one or more layer boundaries; determining, by the computer, a plurality of centerlines corresponding to a plurality of vessels of the choroid layer based on the subset of volumetric data; generating, by the computer, a three-dimensional representation of each vessel of the plurality of vessels in the choroid layer corresponding to each centerline of the plurality of centerlines; and generating, by the computer, a three-dimensional representation of at least a portion of the choroid layer based on the three-dimensional representation of at least a subset of the vessels.
2. The computer-implemented method of claim 1, wherein executing the machine learning model comprises: providing, by the computer, at least a portion of the volumetric data to the machine learning model to cause the machine learning model to generate an output indicating the one or more layer boundaries corresponding to the choroid layer, and wherein determining the subset of the volumetric data comprises: determining, by the computer, the subset of the volumetric data based on the indication of the one or more layer boundaries.
3. The computer-implemented method of claim 2, wherein executing the machine learning model comprises: providing, by the computer, at least a portion of the volumetric data to a first encoder of a plurality of encoders along a contracting path of the machine learning model to cause the machine learning model to generate the output.
4. The computer-implemented method of claim 1, wherein determining the plurality of centerlines corresponding to the plurality of vessels comprises: executing, by the computer, a density -based clustering process to segment the plurality of vessels associated with the subset of volumetric data.
5. The computer-implemented method of claim 1, wherein generating the three-dimensional representation of at least a portion of the choroid layer comprises: generating the three-dimensional representation of at least a portion of the choroid layer, where the three-dimensional representation comprises a choroid thickness map.
6. The computer-implemented method of claim 1, wherein generating the three-dimensional representation of at least a portion of the choroid layer comprises: generating the three-dimensional representation of at least a portion of the choroid layer, where the three-dimensional representation comprises a choroid vascularity index map.
7. The computer-implemented method of claim 1, wherein generating the three-dimensional representation of at least a portion of the choroid layer comprises: generating, by the computer, a heatmap of at least a portion of a retina of a patient based on the three-dimensional representation of at least a subset of the vessels, the heatmap indicating a thickness of the choroid layer at a plurality of points along the retina of the patient.
8. A system comprising a computing device that comprises at least one processor configured to: obtain volumetric data representing three-dimensional imagery of an eyeball based on a plurality of cross-sectional images received from an OCT imaging device;execute a machine learning model configured to identify one or more layer boundaries corresponding to a choroid layer of the eyeball based on the plurality of cross-sectional images of the volumetric data; determine a subset of the volumetric data representing three-dimensional imagery of the choroid layer in response to identifying the one or more layer boundaries; determine a plurality of centerlines corresponding to a plurality of vessels of the choroid layer based on the subset of volumetric data; generate a three-dimensional representation of each vessel of the plurality of vessels in the choroid layer corresponding to each centerline of the plurality of centerlines; and generate a three-dimensional representation of at least a portion of the choroid layer based on the three-dimensional representation of at least a subset of the vessels.
9. The system claim 8, wherein the at least one processor configured to execute the machine learning model is configured to: provide at least a portion of the volumetric data to the machine learning model to cause the machine learning model to generate an output indicating the one or more layer boundaries corresponding to the choroid layer, and wherein the at least one processor configured to determine the subset of the volumetric data is configured to: determine the subset of the volumetric data based on the indication of the one or more layer boundaries.
10. The system of claim 9, wherein the at least one processor configured to execute the machine learning model is configured to: providing at least a portion of the volumetric data to a first encoder of a plurality of encoders along a contracting path of the machine learning model to cause the machine learning model to generate the output.
11. The system of claim 8, wherein the at least one processor configured to determine the plurality of centerlines corresponding to the plurality of vessels is configured to:execute a density -based clustering process to segment the plurality of vessels associated with the subset of volumetric data.
12. The system of claim 8, wherein the at least one processor configured to generate the three- dimensional representation of at least a portion of the choroid layer is configured to: generate the three-dimensional representation of at least a portion of the choroid layer, where the three-dimensional representation comprises a choroid thickness map.
13. The system of claim 8, wherein the at least one processor configured to generate the three- dimensional representation of at least a portion of the choroid layer is configured to: generate the three-dimensional representation of at least a portion of the choroid layer, where the three-dimensional representation comprises a choroid vascularity index map.
14. The system of claim 8, wherein the at least one processor configured to generate the three- dimensional representation of at least a portion of the choroid layer is configured to: generate a heatmap of at least a portion of a retina of a patient based on the three- dimensional representation of at least a subset of the vessels, the heatmap indicating a thickness of the choroid layer at a plurality of points along the retina of the patient.
15. A non-transitory computer-readable medium storing instructions thereon that, when executed by one or more processors, cause the one or more processors to: obtain volumetric data representing three-dimensional imagery of an eyeball based on a plurality of cross-sectional images received from an OCT imaging device; execute a machine learning model configured to identify one or more layer boundaries corresponding to a choroid layer of the eyeball based on the plurality of cross-sectional images of the volumetric data; determine a subset of the volumetric data representing three-dimensional imagery of the choroid layer in response to identifying the one or more layer boundaries; determine a plurality of centerlines corresponding to a plurality of vessels of the choroid layer based on the subset of volumetric data;generate a three-dimensional representation of each vessel of the plurality of vessels in the choroid layer corresponding to each centerline of the plurality of centerlines; and generate a three-dimensional representation of at least a portion of the choroid layer based on the three-dimensional representation of at least a subset of the vessels.
16. The non-transitory computer-readable medium of claim 15, wherein the instructions that cause the one or more processors to execute the machine learning model cause the one or more processors to: provide at least a portion of the volumetric data to the machine learning model to cause the machine learning model to generate an output indicating the one or more layer boundaries corresponding to the choroid layer, and wherein the instructions that cause the one or more processors configured to determine the subset of the volumetric data cause the one or more processors to: determine the subset of the volumetric data based on the indication of the one or more layer boundaries.
17. The non-transitory computer-readable medium of claim 16, wherein the instructions that cause the one or more processors to execute the machine learning model cause the one or more processors to: provide at least a portion of the volumetric data to a first encoder of a plurality of encoders along a contracting path of the machine learning model to cause the machine learning model to generate the output.
18. The non-transitory computer-readable medium of claim 15, wherein the instructions that cause the one or more processors to determine the plurality of centerlines corresponding to the plurality of vessels cause the one or more processors to: execute a density-based clustering process to segment the plurality of vessels associated with the subset of volumetric data.
19. The non-transitory computer-readable medium of claim 15, wherein the instructions that cause the one or more processors to generate the three-dimensional representation of at least a portion of the choroid layer cause the one or more processors to: generate the three-dimensional representation of at least a portion of the choroid layer, where the three-dimensional representation comprises a choroid thickness map.
20. The non-transitory computer-readable medium of claim 15, wherein the instructions that cause the one or more processors to generate the three-dimensional representation of at least a portion of the choroid layer cause the one or more processors to: generate the three-dimensional representation of at least a portion of the choroid layer, where the three-dimensional representation comprises a choroid vascularity index map.
Citation Information
Patent Citations
Automated macular pathology diagnosis in threedimensional (3D) spectral domain optical coherence tomography (sd-oct) images
US20120184845A1
Image processing device, image processing method, and program
US20190274538A1
Medical diagnostic apparatus and method for evaluation of pathological conditions using 3D optical coherence tomography data and images
US20230306568A1