Using deep learning to process images of the eye to predict visual acuity

HK40070075BActive Publication Date: 2026-09-04GENENTECH INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
HK62022058588
Authority / Receiving Office
HK · HK
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-03-16
Filing Date
2022-08-17
Publication Date
2026-09-04
Estimated Expiration
2040-07-30

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately predict visual function based on eye anatomy, especially in retinal diseases such as neovascular age-related macular degeneration, where they cannot reliably predict disease severity and changes in vision.

Method used

Using deep learning models and imaging techniques such as optical coherence tomography and color fundus images, visual acuity is predicted through machine learning models, including convolutional neural networks such as ResNet-50 v2, to process eye images and predict best-corrected visual acuity.

Benefits of technology

It enables more accurate prediction of vision changes, supports early identification of vision disease progression trends and treatment decisions, and improves the reliability and accuracy of vision prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The disclosed systems and methods involve using a machine learning model to process input of a subject's eye and predict the subject's current or future vision. The subject can have been diagnosed with age-related macular degeneration. The predicted current or future vision can be used, for example, to facilitate diagnosing the subject (e.g., with a particular type of age-related macular degeneration), to facilitate determining a treatment strategy for the subject, and / or to facilitate designing a clinical study.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross Reference to Related Applications

[0002] This application claims the benefit of and priority to U.S. Provisional Patent Application Nos. 62 / 882,354 (filed August 2, 2019), 62 / 907,014 (filed September 27, 2019), 62 / 940,989 (filed November 27, 2019), and 62 / 990,354 (filed March 16, 2020). Each of these applications is incorporated herein by reference in its entirety for all purposes. BACKGROUND

[0003] Retinal diseases, such as neovascular age-related macular degeneration (neovascular AMD), are characterized by pathophysiological and anatomical changes that interfere with vision and lead to permanent vision loss. Age-related macular degeneration (AMD) is an eye disease that affects vision in people. There are two types of AMD: dry and wet. Dry AMD is more common and is a milder form of AMD. It usually progresses gradually and slowly affects vision. In contrast, wet AMD (also known as neovascular age-related macular degeneration) is a more severe form of AMD. One risk factor for neovascular AMD is age, where neovascular AMD typically affects people over the age of 50. Smoking also increases a person’s chances of developing neovascular AMD by two to five times. Other risk factors for neovascular AMD include obesity, genetics, race, and gender. People with a body mass index (BMI) over 30 are two and a half times more likely to develop neovascular AMD than people with a lower BMI. In addition, a family history of neovascular AMD also increases the likelihood that some people will develop neovascular AMD. Women, white people, and people with light-colored eyes also have a higher risk of developing neovascular AMD. While risk factors can provide information about the likelihood that some people will develop neovascular AMD, there is currently no technology that can reliably predict the severity of the disease and the rate of degeneration.

[0004] Accordingly, care providers often rely on frequent monitoring to determine whether to initiate and / or change a given treatment (e.g., such as intravitreal anti-vascular endothelial growth factor, which can improve vision in some subjects). One method for objectively assessing visual acuity is to determine which characters on an eye chart (e.g., a Snellen eye chart or a LogMAR) a person correctly identifies. The eye chart can include characters of different sizes (e.g., in the form of presenting different sizes on different lines of the eye chart), and vision can be determined based on determining which size of character an observer correctly identifies. This measure of vision is often referred to as “best-corrected visual acuity” when the observer views the eye chart using eyeglasses or contact lenses.

[0005] One reason that monitoring corrected visual acuity or best corrected visual acuity can provide useful information is that various vision diseases and / or vision conditions can cause changes in the eye that cannot be corrected by eyeglasses or contact lenses. For example, excess fluid (e.g., in the macula) can cause blurring of vision that cannot be corrected by eyeglasses or contact lenses. In contrast, many other types of vision degradation that occur naturally due to aging can be corrected by eyeglasses or contact lenses. Accordingly, corrected visual acuity or best corrected visual acuity can serve as an indicator of the presence, progression, and / or staging of disease.

[0006] While various eye diseases (e.g., including AMD) cause changes in anatomy and further cause degradation in visual function, the accurate relationship for predicting current or future vision based on eye anatomy has not been determined. Conventional correlation analysis is limited in its ability to detect new relationships between anatomical and visual parameters due to the need to identify and pre-specify a set of candidate features for analysis, a process limited by the insight of human researchers. Further, conventional approaches to pre-specifying features, such as central subfield thickness (CST) and foveal thickness (CFT), often result in aggregate measures of retinal health that can not have a sufficiently specific relationship to vision outcomes. For example, in the HARBOR trial data, features derived from intraretinal fluid and total retinal thickness were correlated with baseline BCVA with an R2 of 0.21. 2 = 0.21 at baseline BCVA.

[0007] It is therefore desirable to more reliably predict visual function based on anatomy. BRIEF DESCRIPTION OF DRAWINGS

[0008] The present disclosure is described in connection with the appended drawings:

[0009] FIG. 1 . Process for predicting visual acuity of a subject’s eye using a machine learning model.

[0010] FIG. 2 . Process for conducting a clinical study using vision prediction.

[0011] FIG. 3 . Deep learning pipeline. Best corrected visual acuity (BCVA) is predicted by using three-dimensional optical coherence tomography volumes represented as 30 two-dimensional image inputs to a ResNet-50 v2 convolutional neural network.

[0012] Figure 4. Actual best-corrected visual acuity (BCVA) at concurrent visits versus predicted best-corrected visual acuity. Performance of the deep learning algorithm to analyze optical coherence tomography images to predict BCVA at concurrent visits. (A) Study eye mean across all visits. (B) Fellow eye mean across all visits. (C) Both study eye mean and fellow eye mean across all visits.

[0013] Figure 5. Performance of the deep learning algorithm to predict best-corrected visual acuity (BCVA) < 69 letters at concurrent visits according to the associated optical coherence tomography. (A) BCVA < 69 letters. Study eye at random visits. (B) BCVA < 69 letters. Fellow eye at random visits. (C) BCVA < 69 letters. Both study eye and fellow eye at random visits.

[0014] Figure 6. Actual BCVA versus predicted BCVA at month 12. Performance of the deep learning algorithm to analyze baseline optical coherence tomography images to predict BCVA at month 12. (A) Study eye. (B) Fellow eye. (C) Both study eye and fellow eye.

[0015] Figure 7. Performance of the deep learning algorithm to predict best-corrected visual acuity (BCVA) < 69 letters at month 12 according to the baseline optical coherence tomography. (A) Month 12 BCVA < 69 letters. Study eye. (B) Month 12 BCVA < 69 letters. Fellow eye. (C) Month 12 BCVA < 69 letters. Both study eye and fellow eye.

[0016] FIG. 8 . Standard deviation of BCVA in the fellow eye was limited to the standard deviation of the study eye at baseline.

[0017] FIG. 9 . R2in BCVA by variance for the study eye (black) and fellow eye (red) at baseline, month 6, month 12, month 18, and month 24. 2 .

[0018] FIG. 10 . Deep learning pipeline. Predicting best-corrected visual acuity (BCVA) by using color fundus photograph images input into an initial ResNet-v2 convolutional neural network.

[0019] Figure 11. Performance of the deep learning regression model to analyze color fundus photograph images to predict BCVA in the ANCHOR external validation test set. (A) Actual BCVA versus predicted BCVA at 2 meters from the chart. (B) Actual BCVA versus predicted BCVA at 4 meters from the chart.

[0020] Figure 12. Performance of a deep learning classification model analyzing color fundus photograph images to predict BCVA < 69 letters (Snellen equivalent to < 20 / 40) in the ANCHOR external validation test set. (A) Receiver operating characteristic curve at 2 meters from chart. (B) Receiver operating characteristic curve at 4 meters from chart.

[0021] FIG. 13 a network of computing systems that can be configured to perform one or more actions and / or part or all of one or more methods disclosed herein.

[0022] In the drawings, like reference numerals can be used to denote similar components throughout the several views. Additionally, various components of the same type can be distinguished from one another by following the convention of placing a dash and a second label after the primary reference numeral. For example, a first and a second switch can be referred to as 101 and 101-1, respectively. DETAILED DESCRIPTION

[0023] I. SUMMARY

[0024] The description relates to predicting a measure characterizing a subject's visual function (e.g., vision) based on analysis of an image of the subject's eye. The prediction can aid in defining targets during treatment development and selecting a treatment suitable for a given individual. The image of the eye can include, for example, an optical coherence tomography (OCT) image, a color fundus image, or an infrared fundus image. The subject can, but need not, have been diagnosed with an eye disease (e.g., macular degeneration, neovascular age-related macular degeneration, glaucoma, cataractous or diabetic retinopathy).

[0025] The predicted vision can include a measure of vision generated based on how accurately the subject can recognize visual objects (e.g., visual objects having one or more sizes). The visual objects can include alphanumeric characters and / or the vision can correspond to the ability to read one or more alphanumeric characters on a vision chart. The predicted vision can correspond to a measure corresponding to current vision (e.g., characterizing vision at or over a time period at a time point associated with capture of the image of the eye) or future vision (e.g., characterizing vision at or over a future time period relative to a time point at which the image of the eye was captured).

[0026] A trained machine learning model can be used to predict vision (for a healthy, disease-free subject, for a subject with an eye-related disease, or for a subject with an eye-related medical condition). In some cases, the machine learning model can be configured to output a result of a logic-based conditional evaluation. For example, the machine learning model can be configured and trained to predict whether an observer’s current or future vision (e.g., at a particular time) is below a predefined threshold. The machine learning model can include, for example, a deep machine learning model, a convolutional machine learning model, a regression model, and / or a classifier model. Example deep convolutional neural network architectures that can be used include AlexNet, ResNet, VGGNet, Inception, DenseNet, EfficientNet, GoogLeNet, or a variation of any of these models. The machine learning model can include, for example, one or more convolutional layers, one or more pooling layers (e.g., max-pooling or average-pooling), one or more fully connected layers, one or more dropout layers, and / or one or more activation functions (e.g., a sigmoid function, a tanh function, or a rectified linear unit). The machine learning model can use one or more filter(s) of kernel size in the convolutional layers of the model. For example, the kernel size can vary from layer to layer. The machine learning model can be configured and trained to output, for example, a predicted vision metric, a range of vision metrics (e.g., an open or closed range), and / or a result based on evaluating one or more logic conditions (e.g., in order to indicate whether a subject’s vision is at least as good as a particular threshold).

[0027] A trained machine learning model can be defined based on a set of fixed hyperparameters (e.g., defined by a programmer or user) and a set of parameters (e.g., with values learned during training). The set of parameters can include, for example, a set of weights applied by a node (or neuron) to transform a received value. The value received by a given node can include a set of values from a previous layer (or from an input), where the set of values corresponds to a receptive field of the node (e.g., where the receptive field has a size, padding, etc. as defined by one or more hyperparameters).

[0028] In some cases, the machine learning model is trained to receive an input comprising (e.g., partially or in its entirety) data corresponding to a particular time point or a particular time period, and output a result corresponding to that particular time point or that particular time period. For example, the machine learning model can be trained to predict a subject’s current visual acuity or visual acuity at the time the one or more input images were initially obtained. In some cases, the machine learning model is trained to receive an input comprising (e.g., partially or in its entirety) data corresponding to a particular time point or a particular time period, and output a result corresponding to a later time point or a later time period (e.g., at least 1 week from the time the input data was collected, at least 2 weeks from the time the input data was collected, at least 1 month from the time the input data was collected, at least 3 months from the time the input data was collected, at least 6 months from the time the input data was collected, at least 1 year from the time the input data was collected, at least 2 years from the time the input data was collected, or at least 6 years from the time the input data was collected).

[0029] In some cases, a machine learning model is initialized with parameters learned during training to predict a current metric (e.g., corresponding to a time or time period associated with the input data), and the machine learning model is subsequently further trained (e.g., using transfer learning) to predict a metric for a later time point or a later time period (e.g., at least 1 week from the time the input data was collected, at least 2 weeks from the time the input data was collected, at least 1 month from the time the input data was collected, at least 3 months from the time the input data was collected, at least 6 months from the time the input data was collected, at least 1 year from the time the input data was collected, at least 2 years from the time the input data was collected, or at least 6 years from the time the input data was collected). In some cases, the machine learning model is configured to receive an input value indicating a subsequent time point (e.g., 3 months, 6 months, 1 year, etc.) for which a prediction is to be made relative to the time associated with the set of images included in the input. In some cases, one or more machine learning models are separately trained in order to predict outputs associated with different time points or time periods.

[0030] In some cases, the processing pipeline includes a machine learning model for processing one or more input images to generate an intermediate result (e.g., a prediction of visual acuity or best corrected visual acuity associated with a current time point or a time point associated with collection of the input images), and the processing pipeline further includes one or more other machine learning models and / or one or more other post-processing functions configured to perform the following operations: receive the intermediate result (e.g., a predicted visual acuity at one or more first time points), and output another result (e.g., a predicted visual acuity at a second time point after the first time point). For example, another machine learning model and / or post-processing function can include a regression model configured to receive the intermediate result and possibly one or more other variables (e.g., one or more other intermediate results associated with one or more other time points, one or more other subject characteristics, one or more other subject-specific eye anatomy values, one or more variables indicative of a disease or disease state, etc.) to output a final result (e.g., a subject-specific and / or eye-specific visual acuity metric).

[0031] The one or more predicted visual acuities can be used, for example, to provide information for diagnosis, prognosis, treatment selection, treatment recommendation, dose selection, and / or dose recommendation. The one or more predicted visual acuities can be associated with the same subject in the case of an associated time point or time period at which the images of the subject’s eye are collected, and / or in the case of a time point or time period after the time point or time period at which the images of the subject’s eye are collected. In some cases, a proportional or absolute difference between a predicted visual acuity for a future time point and a predicted visual acuity for a current or prior time point is determined, and the proportional or absolute difference is used to provide information for diagnosis, prognosis, clinical trial criteria, treatment selection, treatment recommendation, dose selection, dose recommendation, recommended medical professional follow-up frequency, etc. Diagnosis can include, for example, identifying a disease (e.g., age-related macular degeneration), a type of disease (e.g., neovascular versus atrophic age-related macular degeneration), or a severity of a disease (e.g., severity of age-related macular degeneration). Prognosis can include, for example, identifying a predicted time at which a given visual function will degrade, a predicted time at which a given disease transition will occur, a probability of a given disease transition occurring (e.g., over a period of time or completely), etc. For example, approximately 10-20% of subjects with atrophic AMD will transition to neovascular AMD. A predicted visual acuity associated with a future time can be used to predict a probability or a predicted time at which a subject will transition to neovascular AMD.

[0032] As one example, if the subject has been diagnosed with age-related macular degeneration, the disease can be“neovascular age-related macular degeneration (AMD)” (also referred to as wet AMD) or“atrophic AMD” (also referred to as dry AMD). Approximately 80% to 90% of individuals with AMD have atrophic AMD. Neovascular AMD is more aggressive and accounts for about 90% of cases of severe vision loss AMD. In some cases (e.g., when the data indicates that the subject has not been diagnosed with neovascular or atrophic AMD), the output can predict that the subject can have neovascular AMD (rather than atrophic AMD) if the relative or absolute difference between the predicted future vision and the current or most recent vision is greater than a predefined change threshold, or if the absolute value of the predicted future vision is greater than a predefined current time threshold. In some cases (e.g., when a preliminary diagnosis of atrophic AMD has been indicated, or in the absence of a report of vascular fluid leakage), the output can predict that the subject can have advanced atrophic AMD (rather than early atrophic AMD) if the relative or absolute difference between the predicted future vision and the current or most recent vision is greater than a predefined change threshold, or if the absolute value of the predicted future vision is greater than a predefined current time threshold. In some cases (e.g., when a preliminary diagnosis of atrophic AMD has been indicated, or in the absence of a report of vascular fluid leakage), the output can predict that the subject is at a relative risk of transitioning to neovascular AMD (e.g., within a given time period) if the relative or absolute difference between the predicted future vision and the current or most recent vision is greater than a predefined change threshold, or if the absolute value of the predicted future vision is greater than a predefined current time threshold.

[0033] Most subjects with atrophic AMD do not receive medication, but are monitored to determine whether the condition progresses to neovascular AMD. If the output predicts that the subject can have neovascular AMD, can have advanced atrophic AMD, and / or is at a high risk of transitioning to neovascular AMD, the computing system can recommend and / or the care provider can decide to: change the plan for image-based monitoring of the subject (e.g., such that images of the eye are collected sooner than previously planned) and / or to begin AMD treatment. AMD treatment can include, for example, an anti-vascular endothelial growth factor (anti-VEGF) agent (e.g., ranibizumab, bevacizumab, aflibercept, and conbercept) or faricimab (a bispecific antibody under clinical investigation that can neutralize angiopoietin-2 and VEGF-A).

[0034] Glaucoma is another eye disease that causes impaired vision (e.g., visual acuity). Glaucoma is caused by damage to the optic nerve due to abnormally high pressure within the eye. Risk factors for glaucoma include age, race, family history, and eye injury. People over the age of 60 are at increased risk of developing glaucoma. African Americans are more likely to develop glaucoma, and they are at increased risk of developing glaucoma over the age of 40. Additionally, medical conditions such as diabetes and hypertension increase the risk of developing glaucoma.

[0035] Glaucoma can be diagnosed through a comprehensive eye exam. Visual acuity testing and intraocular pressure measurements can be used to facilitate the diagnosis of glaucoma. Additionally, OCT and color fundus photography can identify changes or abnormalities in the optic nerve that can be indicative of glaucoma.

[0036] While glaucoma develops slowly, it can lead to blindness within 20 years if left untreated. With treatment, blindness can be prevented. In less severe cases, prescription eye drops can be used to cause the eye to produce less fluid. In more severe cases, laser surgery can be used to enlarge the drainage network within the eye. Imaging techniques such as OCT and color fundus photography can be used to monitor the progression of glaucoma over time. The rate of progression can be helpful in selecting a treatment option.

[0037] Other common eye conditions associated with aging that affect vision are cataracts and dry eye. Cataracts are clouding of the eye lens, while dry eye is a lack of sufficient lubrication in the eye. Cataracts can be removed through surgery, while dry eye can be treated with eye drops or other medications. In extreme cases, surgical measures can be taken to treat dry eye. With treatment of both cataracts and dry eye, normal vision or BCVA can be restored.

[0038] Yet another eye condition that can cause vision loss is diabetic retinopathy. Mild non-proliferative abnormalities such as increased vascular permeability through the blood vessel wall (which increases the flow of small molecules or even whole cells through the blood vessel wall) occur during the early stages of the disease. During the later stages, vascular closure and / or new blood vessel growth on the retinal or vitreous posterior surface are often observed. Macular edema can occur at any stage of the disease and includes fluid accumulation in the macular layer due to blood vessel rupture. Macular edema can cause blurred vision and vision loss.

[0039] All diabetics are at risk of developing diabetic retinopathy. Diagnosis can be made by detecting, for example, decreased vision, microaneurysms (e.g., depicted in fundus photographs or optical coherence tomography (OCT) images), dot-and-blot hemorrhages (e.g., depicted in fundus photographs or OCT images), hard exudates (e.g., depicted in fundus photographs), soft exudates (e.g., depicted in fundus photographs), venous dilation (e.g., depicted in fundus photographs or OCT images), retinal thickening (e.g., depicted in fundus photographs or OCT images), leakage or non-perfusion of the retinal and / or choroidal vasculature (e.g., as indicated by use of fundus fluorescein angiography), and the like. Diabetic retinopathy is further correlated with a plurality of pupil size measurements (e.g., observed after pupil dilation), such as decreased baseline pupil diameter, decreased amplitude of pupil constriction, decreased velocity of pupil constriction, and / or decreased velocity of pupil dilation.

[0040] Early diabetic retinopathy is often left untreated. Progressing diabetic retinopathy can be treated using anti-VEGF, vitrectomy (a surgery to remove blood and scar tissue from the eye), panretinal photocoagulation (a laser treatment to constrict blood vessels), and / or scatter photocoagulation (a laser treatment to inhibit leakage of blood and other fluids in the eye).

[0041] The technology disclosed herein can be used to process images of an eye of a subject having an eye disease or eye condition (e.g., age-related macular degeneration, diabetic retinopathy, macular edema, glaucoma, cataracts, or dry eye) to, for example, predict efficacy of a particular treatment, determine a particular treatment to use or recommend, and / or facilitate design or execution of a clinical study of a particular treatment.

[0042] II. Input data, pre-processing, training data

[0043] II. A. Input data

[0044] Data received by a machine learning model (e.g., to be processed to generate new output or to be processed during training) can include one or more images of one or more processed versions of an eye of a subject. The one or more images can be captured using one or more imaging techniques, such as optical coherence tomography (OCT), color fundus photography, fundus autofluorescence, or infrared fundus photography. The one or more imaging techniques can be non-invasive and / or can not require administration of a dye (e.g., intravenous administration or oral administration). In other cases, the one or more imaging techniques include a technique that includes administration of a dye (e.g., as performed for fundus fluorescein angiography).

[0045] In some cases, the set of input data includes a single image. In some cases, the set of input data includes a set of images, where the images in the set correspond to different depths.

[0046] II. A. 1. Optical coherence tomography

[0047] OCT is a non-invasive imaging technique that uses light waves to construct cross-sectional images of the eye. To generate an image of the eye, an interferometer can split a beam of low-coherence near-infrared light to a target tissue and a reference mirror. The backscattered light received from the target tissue and the reference mirror are combined and the interference of the signals is determined. Target tissue regions that reflect back more light will have a higher interference. The reflectivity information, often referred to as an A-scan, contains information about the longitudinal axis (e.g., depth) of the target tissue. The beam can be directed in a linear direction to generate information at many lateral positions. A cross-sectional “B-scan” image can be generated by combining the depth scans at lateral positions (e.g., using 128, 256, or 512 A-scans). The cross-sectional images can be displayed in real-time, thereby speeding up the process of analysis and diagnosis. Three-dimensional images can be constructed to include multiple B-scans.

[0048] OCT can provide high-resolution images of the eye without contacting the eye. OCT allows for the study of different layers of the retina, which can aid in identifying diseases originating from specific layers. There are two areas in the retina of potential importance, including the optic nerve and the macula. The optic nerve carries information from the eye to the brain, while the macular region is an area with dense photoreceptor cells.

[0049] OCT can provide information about the thickness and size of various layers in the eye, cellular organization, and even axon thickness. Thus, OCT can aid in diagnosing and determining methods of treating conditions such as neovascular AMD. OCT images depict or suggest fluid in the retina, subretinal, or subpigment epithelial space (e.g., via a domelike reflective area in the subretinal space, non-gliotic drusen-like retinal pigment epithelial elevation, or drusen-like pigment epithelial detachment and retinal thickening) that can be consistent with the presence of neovascular AMD.

[0050] II. A. 2. Color fundus photography

[0051] Color fundus photography involves the use of a fundus camera or retinal camera to record color images of the interior surfaces of the eye. A fundus camera is a low-power microscope with a camera that can take pictures of the retina, retinal vasculature, optic disc, macula, and fundus of the eye. Light (e.g., white light) emitted from the fundus camera to capture the image passes in and out through the pupil of the eye. In some cases, the pupil of the eye can be dilated to provide a larger area that can be photographed. Color fundus images depict the retina, macula, blood vessels, optic disc, and fundus.

[0052] Drusen (deposits of lipid primarily under the retina) fluoresce and appear as white or yellow in color fundus images. Drusen are commonly detected in older people, but the presence of large drusen and / or a large number of drusen in the macula is often observed in AMD subjects.

[0053] II. B. Pre-processing

[0054] As noted above, the original images can be processed to generate other images having different perspectives and / or dimensions relative to the original images. For example, multiple A-scans can be processed to generate one or more B-scans, and multiple B-scans can be used to generate a C-scan (e.g., which can capture some three-dimensional information, such as depth). In some cases, the images can include three-dimensional images (e.g., generated based on multiple two-dimensional images). In some cases, color fundus photographs can be collected at different imaging angles, which can facilitate generating images that can convey depth information.

[0055] Because the eye has a curvature, the original images can depict curved structures (e.g., curved retinal layers). Flattening techniques can then be employed to flatten the images (e.g., two-dimensional images or three-dimensional images) based on, for example, a priori estimates of a given structure (e.g., retinal pigment epithelium). For example, the images can be filtered (e.g., using a Gaussian filter) to denoise the images, and the most intense pixel in each column of the denoised images can be identified (e.g., which can represent the retinal pigment epithelium surface). The columns can then be adjusted up or down to align the most intense pixels across the columns. As another example, the retinal pigment epithelium surface can be segmented (e.g., via an intensity threshold), and a function (e.g., a spline function) can be fit to the segmented pixels. The columns of the image can then be realigned to flatten the spline function. In one case, the original images can be flattened to a retinal pigment epithelium cell layer segmentation, and the image volume can be cropped to pixels above and below the flattened retinal pigment epithelium cell.

[0056] The preprocessing can include normalizing and / or standardizing intensity values, which can be performed before or after other types of preprocessing (e.g., before or after generating B-scans, generating C-scans, or applying flattening techniques).

[0057] Flattening techniques and / or other preprocessing techniques can be performed to align depictions of particular structures with target locations. For example, pixels that are inferred to correspond to the retinal pigment epithelium surface can be shifted (e.g., during a flattening process) to a designated row or plane, such that the same row or plane corresponds to the same structure across images.

[0058] Some machine learning models can be configured to receive images of a particular size. Thus, preprocessing can be performed to, for example, crop and / or pad the images so that the size of the images meets the particular size. Some machine learning models can be configured to receive images having a particular resolution. Thus, preprocessing can be performed to, for example, downsample or upsample the images.

[0059] II. C. Training data

[0060] The set of training data includes a plurality of training data elements, each training data element associated with a particular eye of a particular subject. Each training data element of the plurality of training data elements can be further associated with a particular time point and / or medical visit. In some cases, a given set of training data corresponds to a particular type of disease. For example, a set of training data can be defined to correspond to a set of subjects, each subject of the set of subjects having AMD, or each subject of the set of subjects having neovascular AMD. In some cases, one or more other constraints are imposed on the set of subjects (e.g., such that all subjects of the set of subjects are within a particular age range, all subjects of the set of subjects are not receiving a medication, etc.).

[0061] Each training data element of the plurality of training data elements can include input data of one or more images of at least a portion of an eye. In some cases, the images to be processed by the machine learning model are preprocessed versions of original images (e.g., that have been preprocessed using preprocessing techniques such as those disclosed in Section II.B). In some cases, the machine learning model includes one or more preprocessing functions to preprocess received images (e.g., including one or more preprocessing functions disclosed in Section II.B).

[0062] Each training data element of the plurality of training data elements can further include a label that includes or otherwise indicates vision associated with a particular subject and associated with a particular time point, which can indicate an extent to which the subject can discern visual stimuli. The vision metric can be determined by determining whether and / or to what extent the particular subject can accurately identify and / or characterize one or more visual stimuli presented to the subject (e.g., the one or more visual stimuli at a particular distance from the subject and / or having one or more particular sizes). The vision metric can be specific to a particular eye of the subject by evaluating responses provided when the particular subject observes the one or more visual stimuli with only the particular eye (e.g., and the other eye is occluded or closed).

[0063] II. C. 1. Types of vision metrics

[0064] The predicted vision metric can include, for example, a numerical metric such as a ratio, a fraction (e.g., relative to a fixed denominator), a real number, an integer, etc. The predicted vision metric can include a vision category and / or a vision limit. For example, a vision scale can include a set of threshold vision (e.g., 20 / 10, 20 / 20, 20 / 25, 20 / 30, 20 / 40, 20 / 50, 20 / 70, 20 / 100, and 20 / 200 or -0.3, -0.2, -0.1, 0, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, and 0.7) each associated with one or more visual stimuli so as to indicate that an observer has at least the associated threshold vision if the observer can accurately identify the visual stimuli (or features thereof). The predicted vision metric can be defined as the “best” of a set corresponding to results from a particular user (e.g., indicating that the observer can accurately identify or characterize stimuli associated with the vision metric but cannot accurately identify or characterize other stimuli in the set associated with higher vision). In some cases (e.g., when using a Snellen scale), a higher vision metric indicates better vision relative to other metrics associated with lower vision metrics. In some cases (e.g., when using a LogMAR scale), a lower vision metric indicates better vision relative to other metrics associated with higher vision metrics. In some cases (e.g., when using a Jaeger scale), values on the scale do not monotonically depend on vision (e.g., as J1+ indicates better vision than J1, which indicates better vision than J2).

[0065] The visual acuity metric can include a value selected from among values represented along a scale, such as a Snellen scale, a LogMAR scale, or a Jaeger scale. The visual acuity metric can include a Snellen fraction, a LogMAR value, or a Jaeger score. The visual acuity metric can include a value determined based on the subject’s ability to correctly identify and / or characterize one or more characters (e.g., when viewed with a particular eye). The one or more characters can include one or more letters, one or more numbers, one or more tumbling Es, one or more Landolt Cs, and / or one or more Lea symbols presented at a particular size and at a particular distance from the subject.

[0066] The visual acuity metric can include a value determined based on the observer’s ability to identify and / or characterize one or more visual stimuli presented on a chart or card positioned at a particular distance (e.g., 2 meters or 4 meters) from the observer. The chart or card can be a physical chart or card (e.g., including paper, plastic, laminate, cardboard, etc.) or a virtual chart or card (e.g., to be presented on a display of an electronic device). The chart or card can include a plurality of rows, each row including one or more characters (e.g., letters, numbers, tumbling Es, Landolt Cs, and / or Lea symbols). Each row can be associated with a character size indicative of a size (e.g., height, width, and / or aspect ratio) of the one or more characters in that row. The size can vary (e.g., monotonically) between rows. The chart or card can include (e.g., have) a number of rows that is (e.g., is at least) 3 rows, 5 rows, 7 rows, 9 rows, or 11 rows; is less than 20 rows, less than 15 rows, less than 13 rows, less than 10 rows, or less than 7 rows; and / or is about 5 rows, about 7 rows, about 9 rows, and / or about 11 rows.

[0067] The chart or card can include (e.g., be) a Snellen chart (e.g., high greater than wide, including multiple rows of letters, each row having more characters and smaller characters than the previous row, with the character size difference between rows varying throughout the chart— varying in terms of both absolute size difference and relative size difference), a modified Snellen chart, a minimum resolution angle log (LogMAR) chart (where the number of letters per row is the same, the letter size and spacing between rows varies logarithmically between rows, and the letter size is set to be square), a Bailey-Lovie chart, an Early Treatment Diabetic Retinopathy Study (ETDRS) chart (where the number of letters per row is the same, the letter size and spacing between rows varies logarithmically between rows, and the letter size is set to be rectangular), a Rosennbaum card.

[0068] The visual acuity metric can include visual acuity at a particular distance (e.g., 2 meters or 4 meters) with or without vision correction (e.g., glasses or contact lenses). When the visual acuity metric is determined while the observer is using vision correction, the metric can include corrected visual acuity or best-corrected visual acuity.

[0069] The visual acuity metric can include a ratio. The ratio can include a numerator that represents an estimated lower threshold distance (in a particular unit) at which an observer must be relative to a visual stimulus to interpret the stimulus similarly to what an unimpaired observer can interpret at a distance (in a particular unit) as indicated by the denominator of the ratio. The ratio can include a numerator that identifies a maximum distance at which a subject can accurately identify or characterize a visual stimulus (e.g., using a single eye) from the visual stimulus, and the ratio can include a denominator that identifies a maximum distance at which a representative subject (without visual impairment) can accurately identify or characterize the visual stimulus (or a similar visual stimulus having the same or substantially similar height and / or width as the visual stimulus). The ratio can include a Snellen fraction. Ratios and / or tests can be defined to generate ratios having the same numerator value or denominator value. As one example, a 20 / 40 metric can indicate that an observer must be within 20 units of measurement (e.g., feet, yards, meters, etc.) of a visual stimulus to detect and / or interpret a stimulus feature that a typical unimpaired observer can detect and / or interpret at 40 units of measurement. The ratio can include a specified numerator. For example, a model can be configured to output a ratio that includes a ratio with a numerator set to 20.

[0070] In some cases, all of the training data elements include the same type of visual acuity metric. For example, all of the training data elements can include a Snellen fraction. In some cases, at least two of the training data elements include different types of visual acuity metrics. For example, a first training data element can include a Snellen fraction, while a second training data element can include a LogMAR score. In these cases, each of at least some of the training data elements can be processed to convert the visual acuity metric. For example, a lookup table or algorithm can be used to convert a LogMAR score to a Snellen fraction, or a lookup table or algorithm can be used to convert each LogMAR score and each Snellen fraction to a score on yet another scale.

[0071] II. C. 2. Timing associations for vision metrics

[0072] In some cases, for each of the one or more training data elements, the vision metric comprises a vision metric determined based on a functional assessment of the subject’s eye (e.g., via answers provided in response to presentation of a chart or card) that was performed at a time corresponding to the day on which the image in the training data element was collected. For example, the subject can complete the vision assessment (e.g., by attempting to correctly identify or characterize letters or other visual stimuli on an eye chart) on the same day, within the same week, or within the same 2-week period of time as the image of the subject’s eye was collected.

[0073] In some cases, for each of the one or more training data elements, the vision metric comprises a vision metric determined based on a functional assessment of the subject’s eye that was performed at a time after the day on which the image in the training data element was collected. For example, the subject can complete the vision assessment at least or about 1 month, at least or about 2 months, at least or about 6 months, at least or about 1 year, at least or about 2 years, or at least or about 5 years after the day or period of time on which the image of the subject’s eye was collected.

[0074] In some cases, for each of the one or more training data elements, the training data element comprises a vision metric associated with a time corresponding to the day on which the image in the training data element was collected, and the training data element further comprises another vision metric associated with a time after the day on which the image in the training data element was collected.

[0075] III. Machine learning model

[0076] The training data can be used to train a machine learning model, thereby learning a set of parameters. The trained machine learning model can then be used to process other inputs (e.g., of the type described in Section II.A, which can be pre-processed using a type such as disclosed in Section II.B) to generate an output of predicted vision.

[0077] III. A. Model architecture

[0078] The machine learning model can have, for example, a deep learning architecture, a residual network architecture, and / or a convolutional neural network architecture. The deep learning architecture can be configured such that the model performs hierarchical learning to initially identify one or more lower-level low-level features and to identify higher-level higher-level features. The machine learning model can have, for example, a ResNet architecture, an AlexNet architecture, a DenseNet architecture, an EfficientNet architecture, a GoogLeNet architecture, or a VGGNet architecture.

[0079] The machine learning model can include one or more convolutional layers (e.g., at least 3 convolutional layers, at least 4 convolutional layers, at least 5 convolutional layers, or at least 7 convolutional layers), one or more sparse convolutional layers, one or more pooling layers (e.g., max pooling layers or average pooling layers), one or more initial modules, use of dropout, use of batch normalization, one or more dense layers, and / or one or more activation functions. In some cases, each of the one or more or all convolutional layers is followed by a pooling layer. For example, the model can include 5 convolutional layers, and each of the 3 convolutional layers can be followed by a pooling layer. The pooling layers can reduce the number of parameters to be learned by the model.

[0080] The machine learning model can include residual layers or skip connections that feed output from one layer (layer l) to another layer (layer l+2, layer l+3, layer l+4, etc.) that is not immediately adjacent to the one layer. The skip can reduce or avoid vanishing gradient situations. Thus, the skip can further facilitate use of networks with a large number of layers. The machine learning model can include (e.g., in addition to one or more max pooling layers, one or more average pooling layers, and / or one or more fully connected layers) at least 10, at least 20, at least 30, at least 40, or at least 45 convolutional layers. The machine learning model can include residual blocks, each of which includes 3 layers. The size of the kernel used in the early residual blocks of the model can be smaller than the size of the kernel used in the later residual blocks of the model.

[0081] The machine learning model can be configured to output a numerical output or a categorical output (e.g., indicating a level of vision and / or including a binary prediction as to whether the vision is equal to or better than a threshold). In some cases, the machine learning model that outputs a numerical output includes a linear activation function and / or uses an error-based loss function (e.g., mean squared error). The machine learning model that outputs a categorical output can include (e.g.,) a softmax activation function and / or can use a categorical loss function (e.g., a sparse categorical cross-entropy loss function).

[0082] III. B. Model training

[0083] In some cases, the first set of training data includes a first set of training data elements, each including one or more images (of a subject’s eye) and a label corresponding to the subject’s vision at a time corresponding to the time at which the images were captured. The first set of training data can be used to train a machine learning model having a particular architecture, thereby learning a first set of parameters.

[0084] The second set of training data can include a second set of training data elements, each including one or more images (of the subject’s eye) and a label corresponding to the subject’s vision at a time later (e.g., 12 months later) than the time at which the images were captured. In some cases, each subject represented in the first set of training data is also represented in the second set of training data. In some cases, at least some of the subjects represented in the first set of training data are also represented in the second set of training data. In some cases, the subjects represented in the first set of training data are at least partially different or completely different from the subjects represented in the second set of training data.

[0085] A machine learning model having at least partially the same or completely the same architecture as the particular architecture can be trained using the second set of training data, thereby learning a second set of parameters. For example, each model trained using the first and second sets of training data can include a ResNet-50 CNN architecture.

[0086] In some cases, a model to be trained to predict subsequent vision metrics is initialized using parameters learned by a model trained to predict current (e.g., corresponding to the same time period at which the input images were collected) vision. Subsequent training can be performed thereafter.

[0087] In some cases, a single machine learning model can be trained using both the first and second sets of training data, but the machine learning model can be configured to receive an indication (in addition to the images) as to which time point to predict to.

[0088] In some cases, the machine learning model can be configured to predict a rate of vision decline, one or more rate constants, and / or the like, to predict how the subject’s eye’s vision will change over time (e.g., how quickly the vision will decline). The prediction can include constants to be used in (e.g.) a linear function, a logarithmic function, a polynomial function, and / or an exponential function to characterize the vision over time.

[0089] IV. Using model output to provide information for clinical study design, prognosis, and treatment selection

[0090] The model output predicting current vision metrics, predicting vision metrics at particular future time points, and / or predicting how the vision will change can be used, for example, to provide information for clinical trial design, subject prognosis, and selection of treatments.

[0091] IV. A. Clinical study design

[0092] In some cases, a clinical study can be defined such that the eligibility criteria is related to observed or predicted vision. For example, a clinical study can be defined to enroll subjects with vision equal to or worse than 20 / 80 (and these subjects meet other eligibility criteria). The vision-based eligibility criteria can be defined in relation to a vision metric that characterizes current or past vision, and that has been generated based on images of at least a portion of a subject’s eye. The criteria can require a predicted metric, or require a predicted metric or an observed metric. The vision-based eligibility criteria can alternatively or additionally be defined to involve a predicted vision metric that characterizes predicted vision at a subsequent time (e.g., requires predicted vision at a time point that is one year from the current time to be worse than 20 / 40 vision).

[0093] Many clinical studies (e.g., clinical trials) are controlled such that a first group of subjects receives a study treatment, while a second group of subjects receives a different treatment, receives a control treatment, receives no treatment, and / or receives standard of care. Model results corresponding to predicted vision can be used to facilitate selection of groups of subjects such that the population of subjects in the first group is similar to the population of subjects in the second group. Currently, there is a large variation in the rate of decline in visual function among AMD subjects. Thus, two AMD subjects who once had the same vision can have significantly different visual function at a second time. This can lead to unintentionally defining groups of subjects in such a way that one group can progress much faster than the other group, which can confound the results.

[0094] Accordingly, stratification of a clinical study can be designed to include eligibility criteria that specifies a range of predicted vision at a future time point. For example, a clinical study can require that OCT images or color fundus images be collected over a first time period (prior to the start of the trial), and that the predicted vision at a second time point (e.g., 12 or 18 months from the first time period) be within a pre-defined range. The criteria can further identify a constraint (e.g., a range) on the predicted vision corresponding to the time at which the images were collected.

[0095] Additionally or alternatively, the predicted vision can be used to process clinical trial data and develop indications and / or hypotheses about what types of subjects are likely to particularly benefit (or alternatively, are particularly unlikely to benefit) from a given treatment. For example, during or after a trial, for each subject in a treatment group, a current metric can be determined, such as a current vision (e.g., as determined using a vision chart or card), a predicted current vision (e.g., as determined using a current image of the eye), a current diagnosis of AMD (e.g., whether the AMD is wet or dry), etc. It can then be determined whether a predicted vision metric corresponding to a pre-trial time (e.g., associated with a time period in which the images were acquired) or another predicted vision metric corresponding to a subsequent time is a prediction of the current metric.

[0096] IV. B. Prognosis

[0097] The predicted vision metric can be used as part of or to inform a prognosis of the subject. A subject can have a worse prognosis when the predicted vision of the subject at a subsequent time (e.g., within 12 months) is worse than a threshold or corresponds to a more significant deterioration than a subject whose predicted subsequent vision is better than the threshold or corresponds to a less significant deterioration. The prognosis can affect the prospects of the subject, treatment selection, etc.

[0098] IV. C. Treatment selection

[0099] There are significant differences in the degree to which different subjects respond to a given treatment. For example, anti-vascular endothelial growth factor (anti-VEGF) is the standard of care for neovascular AMD. However, there is a high degree of variability between subjects in terms of the degree to which they respond.

[0100] There can be disease activity when, for example, there is a decrease or no improvement in vision, there is no decrease in central retinal thickness observed by OCT, new intraretinal fluid is detected, new subretinal fluid is detected, and / or new retinal thickening is detected. Thus, the predicted future vision can inform whether a given subject will respond effectively to a particular treatment. For example, a model can be trained using a training dataset that includes images of a subject’s eye prior to administration of a treatment and that includes (as a label) a predicted vision (or a predicted change in vision) at a subsequent time while the subject received a particular treatment between a first time period (associated with the image collection) and a second time point (associated with the subsequent vision).

[0101] Accordingly, the care provider can be able to predict the extent to which vision will change when using a particular treatment. Even multiple models can be used to evaluate different potential treatment options. The care provider can then be able to predict whether a particular treatment will be effective for a particular subject and / or which of a plurality of treatment options will be most effective for a particular subject.

[0102] For example, for neovascular AMD, one treatment option is anti-VEGF injections (where VEGF stands for a protein called vascular endothelial growth factor, which can promote the formation of new blood vessels in the back of the eye, which can cause macular degeneration due to leakage of blood and other fluids). Anti-VEGF injections are injections into the vitreous of the eye to stop the abnormal growth of blood vessels. The most common anti-VEGF injections are ranibizumab, bevacizumab, conbercept, and aflibercept. A person’s response to anti-VEGF injections is often monitored by manual assessment of OCT or color fundus photography to determine how frequently injections should be received.

[0103] Anti-VEGF injections can be received monthly until it is determined that the injection interval can be lengthened. Vision loss due to neovascular AMD is often (but not always) stopped by the use of anti-VEGF injections. Sometimes, a subject will recover some of the lost vision from anti-VEGF injections.

[0104] Another treatment option (currently in phase III studies) is a port delivery system (PDS) used in conjunction with anti-VEGF drugs. The PDS is a permanent refillable ocular implant that can continuously deliver anti-VEGF drugs to the eye for months. The PDS can allow a subject to see an ophthalmologist only twice a year to refill the drug. Reducing the number of visits to an ophthalmologist needed can reduce the burden of treatment that often leads to inadequate treatment and poor vision outcomes.

[0105] Yet another treatment option is faricimab, an antibody that is in phase III studies. Faricimab is administered intraocularly. Some of the lost vision can be recovered from this treatment.

[0106] Further, yet another treatment option for neovascular AMD is photodynamic therapy (PDT). PDT is a laser treatment that can break down excess blood vessels. If a subject experiences vision loss that is gradual rather than sudden, a doctor can choose PDT over anti-VEGF injections. In some cases, PDT can be combined with anti-VEGF therapy to slow down the damage to central vision caused by neovascular AMD.

[0107] Accordingly, model predictions can be used to provide information for deciding whether to recommend or use, for a particular subject and / or in the eye of a subject, an anti-VEGF agent, an anti-VEGF agent delivered via intraocular injection, a particular VEGF agent (e.g., ranibizumab or bevacizumab), PDS, or PDT.

[0108] IV. D. Treatment initiation and / or monitoring changes

[0109] In some cases, machine learning model outputs can be used to recommend, prescribe, or administer an AMD treatment to a subject who is not at that time receiving an AMD treatment or who can not be receiving any AMD treatment. For example, a subject can have been diagnosed with atrophic AMD, which is generally not treated with medication. However, an output of a machine learning model (e.g., trained using data corresponding to subjects diagnosed with atrophic AMD, subjects diagnosed with neovascular AMD, or both) can correspond to a prediction that visual acuity (e.g., best-corrected visual acuity) corresponding to a baseline day (during which the input image was collected), corresponding to a concurrent visit, or corresponding to a subsequent time (e.g., 3 days, 5 days, 1 week, 2 weeks, 1 month, 6 months, 1 year, 2 years, 5 years from the baseline day) will be worse than a predefined threshold (e.g., corresponding to a Snellen fraction of 20 / 30, 20 / 40, 20 / 60, or 20 / 80). Alternatively or additionally, an output from a machine learning model (e.g., trained using data corresponding to subjects diagnosed with atrophic AMD, subjects diagnosed with neovascular AMD, or both) can correspond to a prediction that visual acuity will decrease by at least a threshold number of letters or an absolute predefined amount (e.g., corresponding to a Snellen fraction at a subsequent time that is less than 90%, 75%, 66%, or 50% of a Snellen fraction at a baseline time).

[0110] In response to one or more of these types of predictions, a computing system (e.g., a computing system that generated the predictions, or a computing system that received the predictions from another computing system) can output a recommendation of a particular action that a care provider (e.g., a doctor, a nurse, a doctor’s office, a hospital, etc.) can provide and / or can perform. The particular action can include, for example: changing a diagnosis for the subject (e.g., changing from atrophic AMD to neovascular AMD, or changing from early atrophic AMD to late atrophic AMD), initiating an AMD treatment for the subject (e.g., including one or more treatment regimens identified herein), changing a treatment plan for the subject (e.g., increasing a frequency of anti-VEGF dosing), and / or changing an AMD treatment regimen for the subject (e.g., to a treatment regimen identified herein). For example, the particular action can include initiating an anti-VEGF treatment for the subject. In some cases, the particular action includes changing a monitoring plan such that, for example, a next date for performing an imaging of the subject’s eye or a vision test is scheduled or changed (e.g., to an earlier date), a frequency of performing imaging of the subject’s eye is increased, and / or a frequency of testing the subject’s vision is increased.

[0111] V. Exemplary process

[0112] FIG. 1 A process 100 for predicting vision of a subject’s eye using a machine learning model is depicted. The process 100 begins at block 105, in which a machine learning model is trained using a set of training images and a set of labels to learn a set of parameters. Each training image in the set of training images can depict at least a portion of an eye of a training subject, such that the set of training images depicts at least a portion of each eye of each training subject in the set of training subjects. The set of training images can include, for example, one or more color fundus photographs, one or more OCT images, and / or one or more other images. Each label in the set of labels can include, for example, a measure of vision, such as a numerical visual acuity, an identification of a range of vision, and / or an indication of whether the training subject’s eye has a vision below a particular threshold. The measure of vision can characterize the vision of the training subject’s eye at a time at which the training image was collected (e.g., on the same day as the image was collected) or at a time after the training image was collected (e.g., about 6 months, about 12 months, about 2 years, or about 5 years after the date on which the training image was collected).

[0113] In some cases, each training subject in the set of training subjects has been diagnosed with a vision disease or vision medical condition, such as macular degeneration, age-related macular degeneration, neovascular macular degeneration, glaucoma, diabetic retinopathy, cataract, or macular edema (e.g., prior to collecting their training images). In some cases, each of one or more training subjects or each of all of the training subjects in the set of training subjects has not been diagnosed with a vision disease or vision medical condition. In some cases, the set of training subjects includes both training subjects with a vision disease or vision condition and training subjects without a vision disease or vision condition.

[0114] The machine learning model can include a model disclosed herein, such as a deep convolutional neural network. The machine learning model can include one or more residual connections and / or one or more feed-forward layers. The set of parameters can include a set of weights and / or one or more convolutional kernels.

[0115] The machine learning model can further include one or more pre-processing functions that can pre-process an image before sending the pre-processed image to, for example, a neural network. The machine learning model can further include one or more post-processing functions that can, for example, convert an output from a neural network to another result. For example, a post-processing function can convert a result to a particular visual acuity scale, or can identify a range in which a numerical neural network output falls.

[0116] At block 110, an image of at least a portion of an eye of a particular subject is accessed. The particular subject can be different from each of the set of training subjects and / or can have been one of the set of training subjects. The image accessed at block 110 can include a fundus photograph, a color fundus photograph, an OCT image, or another image. The image can include a digital image. The at least a portion of the eye can include, for example, at least a portion of: a retina, a macula, an optic disc, a lens, a pupil, and / or an iris. For example, the image can depict at least a portion of a retina and at least a portion of a macula.

[0117] At block 115, the image of at least a portion of the eye of the particular subject is input into the trained machine learning model. The trained machine learning model can determine, based on the image, a vision metric corresponding to a predicted vision of the eye. The predicted vision can include, for example, a numerical vision, a categorical vision, a vision range, and / or an indication as to whether the predicted vision exceeds a threshold.

[0118] At block 120, the vision metric is returned. For example, the vision metric can be presented (e.g., displayed) via a user interface. As another example, the vision metric can be transmitted to another device (e.g., from a server to a user device).

[0119] FIG. 2 A process 200 is depicted that uses vision prediction for clinical studies. At block 205, the vision of each subject in a set of subjects is predicted. The vision can include the vision of the subject’s eye. The predicted vision can include a vision metric predicted based on an image of at least a portion of the subject’s eye. The vision can be predicted using the methods herein, such as all of the portions of process 100. The predicted vision can be a vision predicted for a time corresponding to the capture of the image or a subsequent time. The predicted vision can be a vision predicted for a particular time corresponding to a clinical study (e.g., 6 months after treatment begins, 12 months after monitoring begins, etc.).

[0120] At block 210, for each subject in the set of subjects, it can be determined whether the subject is eligible to participate in a particular clinical study. The determination can be based on whether the subject satisfies each of a set of eligibility criteria. A given criterion can include logic such that (for example) the given criterion is satisfied if any of a plurality of conditions is satisfied.

[0121] The set of criteria can indicate that the subject should be within a particular age range, should be diagnosed with a particular vision disease or condition (e.g., age-related macular degeneration, neovascular age-related macular degeneration), should not have one or more particular other diseases, etc. In some cases, the set of subjects includes subjects enrolled in the clinical study.

[0122] The set of criteria can include a criterion that uses a predicted vision of the subject. For example, the criterion can indicate that the subject should be associated with a predicted vision metric that is above a predefined threshold. As another example, the criterion can indicate that the subject should be associated with a change in vision metric that is above a predefined threshold. The change in vision metric can include an absolute or relative difference between a vision metric associated with a time point relative to a vision metric associated with a baseline time point. The time point can include, for example: a current time, a time a predefined duration after the start of treatment (e.g., 6 months after the start of treatment, 12 months after the start of treatment), a time a predefined duration after the start of a clinical study, a particular date, etc. The baseline time point can correspond to, for example: a time before the start of a clinical study or treatment of the subject’s eye, a time of the start of a clinical study or treatment of the subject’s eye, etc. In some cases, each of the vision metric associated with the time point and the vision metric associated with the time point and the vision metric associated with the baseline time point includes a vision metric generated based on processing (e.g., using a machine learning model) of at least a portion of an image of the subject’s eye. For example, an image of at least a portion of the subject’s eye collected at the baseline time point can be used to generate a predicted vision metric of the subject’s eye at the baseline time point as well as a vision metric of the subject’s eye at a subsequent time point. As another example, an image of at least a portion of the subject’s eye can be collected at each of the baseline time point and the subsequent time point, and the image can be processed to generate a respective predicted vision metric. In some cases, one of the vision metric associated with the time point and the vision metric associated with the time point and the vision metric associated with the baseline time point is generated based on a vision test (e.g., using a vision chart or a vision chart card), and the other of the vision metric associated with the time point and the vision metric associated with the time point and the vision metric associated with the baseline time point is generated by processing at least a portion of an image of the subject’s eye.

[0123] At block 215, a clinical study is conducted with a subset of the set of subjects, where each subject in the subset is determined to be eligible to participate in the clinical study. For each subject in the subset, it can be determined that the criteria related to vision are satisfied. Conducting the clinical study can include, for example, dividing the subset of subjects into two groups and providing a study treatment to the subjects in one group. The subjects in the other group can receive, for example, a different treatment, no treatment, a different dosage of the study treatment, the study treatment in a different type of administration or formulation, standard of care, etc.

[0124] At block 220, results of the clinical study are generated. For example, the results can indicate the degree of effectiveness of the investigational treatment in slowing, halting, or reversing a given medical condition (e.g., as indicated via vision tests, eye imaging, or other tests). Efficacy can be assessed by comparing vision metrics or other medical metrics between the first subset and the second subset, between subjects who received a given treatment and subjects who did not receive the given treatment, etc. In some cases, the results indicate the degree of observed vision of a subject after a period of treatment as compared to the subject’s vision based on processing of images of at least a portion of the subject’s eye collected at a baseline time.

[0125] VI. Examples

[0126] VI. A. Machine learning model for processing OCT images to predict vision metrics

[0127] This example assessed whether deep learning could automatically predict concurrent and future BCVA from OCT images from subjects with neovascular AMD in the Phase 3 HARBOR clinical study (NCT00891735, referred to herein as “HARBOR”). Specifically, the results were related to: (1) a model that assessed the quality of deep learning models to accurately predict best-corrected visual acuity (BCVA) values from OCT images, and (2) models that predicted BCVA < 69 letters (Snellen equivalent, 20 / 40), < 59 letters (Snellen equivalent, 20 / 60), or < 38 letters (Snellen equivalent, 20 / 200) from OCT images. For BCVA outcomes, the ability of a deep learning model to predict BCVA from OCT images taken at the same (concurrent) visit was assessed, as well as the ability of that deep learning model to predict 12-month BCVA from baseline OCT images. The Snellen equivalents of 20 / 40 and 20 / 60 were chosen because vision worse than these levels is considered to reflect impaired vision according to the U.S. and World Health Organization definitions, respectively. A Snellen equivalent of 20 / 200 or worse was used to reflect the U.S. definition of legal blindness.

[0128] VI. A. 1. Methods

[0129] VI. A. 1. a. Data sources

[0130] BCVA measurements and OCT images taken prospectively from 1071 subjects in the Phase 3 HARBOR clinical study were used. HARBOR adhered to the principles of the Declaration of Helsinki and was in compliance with the Health Insurance Portability and Accountability Act. The protocol was approved by each institutional review board prior to study initiation, and written informed consent was provided for all subjects regarding future medical research and analyses based on trial results.

[0131] The HARBOR study enrolled 1097 adult subjects with newly treated central subfoveal CNV secondary to neovascular AMD, who had a BCVA between 20 / 40 and 20 / 320 (Snellen equivalent) using standard ETDRS charts and protocol. Subjects (one study eye per subject) were randomized in a 1 : 1 : 1 : 1 ratio to receive ranibizumab: 0.5 mg monthly, 0.5 mg as needed (PRN), 2.0 mg monthly, and 2.0 mg PRN, according to the following treatment regimen. Subjects in the PRN groups received three monthly injections, followed by monthly assessments, with retreatment only if there was any evidence of disease activity on OCT images or a decrease in BCVA of >5 letters from the previous visit. BCVA measurements and OCT images were obtained at baseline and monthly over the 24-month period.

[0132] OCT images. Images were collected using a spectral domain Cirrus HD-OCT IMAGE instrument (Carl Zeiss Meditec, Dublin, CA, USA). The resolution was 200 x 200 x 1024 voxels, the size was 30.0 x 30.0 x 2.0 pm, and the covered volume was 6 x 6 x 2 mm. The dataset consisted of 50,275 OCT image scans from 1071 subjects. For each of the 50,275 OCT IMAGE scans, the retina was flattened to the retinal pigment epithelium layer segmentation provided by the Zeiss software, and the volume was cropped to 384 pixels above and 128 pixels below the flattened retinal pigment epithelium. Thirty 512 x 200 pixel slices were generated per scan by rotating the center of the cropped volume around the z-axis at angles of 0, 30, 60, 90, 120, and 150 degrees (offset at -8, -4, 0, 4, and 8 pixels for each of the six angles), resulting in a total of 1,508,250 slices. The OCT IMAGE dataset was split at the subject level into (1) a randomly selected internal validation test set of 147 subjects to be used for evaluation (Table 1), and (2) a set of 924 subjects to be used for model development via cross-validation, which was further split into five folds (Table 2). For each outcome variable, the subjects in each fold were kept constant.

[0133]

[0134]

[0135] BCVA, best-corrected visual acuity; OCT image, optical coherence tomography; SD, standard deviation.

[0136] BCVA from screening to month 24.

[0137] Table 1. Characteristics of the internal validation test set used to evaluate the model to predict BCVA from OCT images.

[0138]

[0139]

[0140] BCVA, best-corrected visual acuity; OCT image, optical coherence tomography; SD, standard deviation.

[0141] BCVA from screening to month 24.

[0142] Table 2. Characteristics of the OCT image dataset used to develop the deep learning model to predict BCVA from OCT images.

[0143] VI. A. 1. b. Outcome variables for deep learning modeling

[0144] BCVA. BCVA outcomes of interest were (1) BCVA in ETDRS letters at each study visit, and (2) whether a particular BCVA value was < 69 letters (Snellen equivalent, 20 / 40), < 59 letters (Snellen equivalent, 20 / 60), or < 38 letters (Snellen equivalent, 20 / 200). The Snellen equivalents of 20 / 40 and 20 / 60 were chosen because they were considered to reflect functionally meaningful levels of vision impairment, and the Snellen equivalent of 20 / 200 or worse was used to reflect the U.S. definition of legal blindness.

[0145] Mean (± standard deviation [SD]) BCVA for study eyes in the internal validation test set was 53.93 (± 13.20) letters at baseline and 65.02 (± 17.12) letters at month 24; mean (± SD) BCVA for fellow eyes was 69.46 (± 22.92) letters at baseline and 68.56 (± 22.56) letters at month 24 (Table 1). The range of vision for study eyes at baseline, month 6, month 12, month 18, and month 24 was 55, 83, 89, 82, and 78 letters, respectively (Table 1). The range of vision for fellow eyes at baseline, month 6, month 12, month 18, and month 24 was 93, 98, 100, 96, and 98 letters, respectively (Table 1).

[0146] VI. A. 1. c. Deep learning algorithm

[0147] The ability of the DL models to predict the following was evaluated: (1) accurate BCVA values in ETDRS letters from OCT images obtained in the same visit; (2) accurate BCVA at month 12 from baseline OCT images; (3) BCVA < 69 letters (Snellen equivalent, 20 / 40), < 59 letters (Snellen equivalent, 20 / 60), or < 38 letters (Snellen equivalent, 20 / 200) from OCT images obtained at the same visit; and (4) BCVA < 69 letters, < 59 letters, or < 38 letters at month 12 from baseline OCT images.

[0148] BCVA at concurrent visits was predicted. Deep learning modeling was performed using TensorFlow (1.14.0) and Keras (2.2.5) on an Nvidia V100 GPU and a ResNet-50 v2 CNN architecture. Individual slices of 512 x 200 pixels were randomly shuffled from the training set and input into the CNN using a batch size of 64 images. The model was trained to predict BCVA in the same visit as the OCT image scan FIG. 3 ) In some cases, L2 regularization (0.05) was used on the final dense layer with a linear activation function, and a layer of global average pooling and dropout (0.85) was added to the CNN. This model is referred to as the “concurrent visit regression model” in this example. The loss function was mean squared error, and the optimizer was RAdam. For each cross-validation fold, the model was trained for only one epoch.

[0149] In some cases, the concurrent visit regression model architecture was used except that the last layer used a softmax activation function with a sparse categorical cross-entropy loss function. The model was initialized with the weights from the regression model from each fold. The model was trained for two epochs using the Adam optimizer with the base model layers that could not be trained, and then an additional epoch was trained using the stochastic gradient descent (SGD) optimizer with the base model layers that could be trained.

[0150] BCVA at 12 months from baseline was predicted. For the regression task of predicting BCVA at month 12 from baseline OCT images, deep learning modeling was performed using TensorFlow (1.14.0) and Keras (2.2.5) on an Nvidia V100 GPU and a ResNet-50 v2 CNN architecture. Individual slices of 512 x 200 pixels were randomly shuffled from the training set and input into the CNN using a batch size of 64 images. The model was trained to predict BCVA in the same visit as the OCT image scan FIG. 3). In some cases, global average pooling and dropout (0.995) layers were added to the CNN with L2 regularization (0.05) on the final dense layer using a linear activation function. For each fold, the model was initialized with the weights from a regression model trained to predict BCVA at the concurrent visit. The first three epochs were trained using the SGD optimizer with the non-trainable base model layers, then an additional 1000 epochs were trained using the SGD optimizer with the trainable base model layers. This model is referred to in the present example as the “12-month regression model”.

[0151] For classification, the 12-month regression model architecture was used except that the last layer used a softmax activation function with a sparse categorical cross-entropy loss function. For each fold, the model was initialized with the weights from a regression model trained to predict BCVA at 12 months from baseline. The model was trained for 20 epochs using the RAdam optimizer with the non-trainable base model layers. The weights used for prediction were selected from the epoch with the lowest validation loss for each fold.

[0152] Evaluation of deep learning models. The metric used to evaluate model fit at a particular visit was calculated on an eye level by averaging the predictions for each eye generated for each of the five development models on the 30 slices for each eye from the out-of-sample internal validation test set. In other words, each of the five-fold cross-validation models that have seen 80% of the data from the training set were used to generate a prediction for each of the 30 slices for each eye in the test set, resulting in 150 predictions per eye, of which the mean was taken. Furthermore, to evaluate model performance in concurrent visits across all visits, while accounting for the potential bias of repeated measures for the same eye, the mean of all visits was determined for the regression task and a random visit was selected for each subject in the classification task. R 2Root mean square error (RMSE) and mean difference (MD) and 95% limits of agreement (LOA) were used to evaluate the deep learning regression models, while area under the receiver operating characteristic curve (AUC) and area under the precision recall curve (AUPRC) were used to assess the classification performance of the deep learning models. To understand whether the deep learning prediction of 12-month BCVA from baseline OCT images provided additional information compared to baseline BCVA alone, linear models were fitted using the R statistical programming language to predict 12-month BCVA from: (i) the univariate input of deep learning prediction of 12-month BCVA from baseline OCT images, (ii) the univariate input of baseline BCVA, and (iii) the multivariate input of both deep learning prediction of 12-month BCVA from baseline OCT images and baseline BCVA. In addition, the mean of the 30 times adjusted predictions per eye per visit were used to report the results of the five-fold cross-validation adjusted set (Tables 3-6).

[0153] VI. A. 2. Results

[0154] VI. A. 2. a. Predicting best corrected visual acuity (BCVA) at concurrent visits

[0155] Regression results. In the study eyes, the deep learning model to predict BCVA at concurrent visits had R 2 = 0.24, RMSE = 11.55, MD = -1.81 letters, 95% LOA, -26.57 to 22.95 letters, and a mean of R 2 = 0.67, RMSE = 8.60, MD = 0.04 letters, 95% LOA, -16.96 to 17.04 letters for all visits (Table 3; FIG. 4A ). In the fellow eyes, at baseline, R 2 = 0.80, RMSE = 10.35, MD = -1.86 letters, 95% LOA, -23.03 to 19.31 letters, and a mean of R 2 = 0.84, RMSE = 9.01, MD = 0.51 letters, 95% LOA, -17.58 to 18.61 letters for all visits (Table 3; FIG. 4B ). In all eyes, at baseline, R 2 = 0.66, RMSE = 11.75, MD = -1.84 letters, 95% LOA, -24.83 to 21.16 letters, and a mean of R 2 = 0.79, RMSE = 8.78, MD = 0.28 letters, 95% LOA, -17.26 to 17.81 letters for all visits (Table 3; FIG. 4C ).

[0156]

[0157] BCVA, best-corrected visual acuity; OCT image, optical coherence tomography; CI, 95% confidence interval; RMSE, root mean square error.

[0158] Table 3. Performance of deep learning models for regression of BCVA on OCT images

[0159]

[0160] AUC, area under the receiver operating characteristic curve; BCVA, best-corrected visual acuity; OCT image, optical coherence tomography; CI, 95% confidence interval.

[0161] Table 4. Performance of deep learning models for binary classification of BCVA < 69 letters (Snellen equivalent, 20 / 40), < 59 letters (Snellen equivalent, 20 / 60), and < 38 letters (Snellen equivalent, 20 / 200) at 12 months concurrent with associated optical coherence tomography visits.

[0162]

[0163] BCVA, best-corrected visual acuity, OCT, optical coherence tomography, CI, 95% confidence interval, RMSE, root mean square error; *, P < 0.001 for single-intercept model compared with null hypothesis; **, P < 0.001 for each coefficient in multivariate model.

[0164] Table 5. Cross-validation results. Performance of linear models for regression of BCVA from baseline OCT and baseline letters at 12 months.

[0165]

[0166] AUC, area under the receiver operating characteristic curve; BCVA, best-corrected visual acuity; OCT, optical coherence tomography; CI, 95% confidence interval.

[0167] Table 6. Cross-validation results. Performance of deep learning models for binary classification of BCVA < 69 letters (Snellen equivalent, 20 / 40), < 59 letters (Snellen equivalent, 20 / 60), and < 38 letters (Snellen equivalent, 20 / 200) at 12 months from baseline OCT.

[0168] Classification model to predict BCVA < 69 letters (Snellen equivalent, 20 / 40). In study eyes, the deep learning model to predict BCVA < 69 letters (Snellen equivalent, 20 / 40) had an AUC = 0.89 and an AUPRC = 0.88 for a concurrent visit at random for each eye with a class balance of 72 positive eyes and 75 negative eyes. (Table 4; FIG. 5A ). In fellow eyes, the deep learning model to predict BCVA < 69 letters (Snellen equivalent, 20 / 40) had an AUC = 0.93 and an AUPRC = 0.97 for a concurrent visit at random for each eye with a class balance of 103 positive eyes and 44 negative eyes. (Table 4; FIG. 5B ). In all eyes, the deep learning model to predict BCVA < 69 letters (Snellen equivalent, 20 / 40) had an AUC = 0.92 and an AUPRC = 0.94 for a concurrent visit at random for each eye with a class balance of 175 positive eyes and 119 negative eyes. (Table 4; FIG. 5C ).

[0169] Classification model to predict BCVA < 59 letters (Snellen equivalent, 20 / 60). In study eyes, the deep learning model to predict BCVA < 59 letters had an AUC = 0.92 and an AUPRC = 0.95 for a concurrent visit at random for each eye with a class balance of 100 positive eyes and 47 negative eyes. (Table 4). In fellow eyes, the deep learning model to predict BCVA < 59 letters had an AUC = 0.97 and an AUPRC = 0.99 for a concurrent visit at random for each eye with a class balance of 114 positive eyes and 33 negative eyes. (Table 4). In all eyes, the deep learning model to predict BCVA < 59 letters had an AUC = 0.95 and an AUPRC = 0.98 for a concurrent visit at random for each eye with a class balance of 214 positive eyes and 80 negative eyes. (Table 4).

[0170] Classification models predicting BCVA < 38 letters (Snellen equivalent, 20 / 200). In study eyes, the deep learning model for predicting BCVA < 38 letters had an AUC = 0.92 and an AUPRC = 0.99 for a single concurrent visit at random for each eye, with a class balance of 113 positive eyes and 14 negative eyes. (Table 4). In fellow eyes, the deep learning model for predicting BCVA < 38 letters had an AUC = 0.98 and an AUPRC = 1.00 for a single concurrent visit at random for each eye, with a class balance of 129 positive eyes and 18 negative eyes. (Table 4). In all eyes, the deep learning model for predicting BCVA < 38 letters had an AUC = 0.96 and an AUPRC = 0.99 for a single concurrent visit at random for each eye, with a class balance of 262 positive eyes and 32 negative eyes. (Table 4).

[0171] VI. A. 2. b. Predicting BCVA at 12 months from baseline OCT images

[0172] Regression results. Characteristics of the dataset used to assess the ability of the model to predict BCVA at 12 months from the baseline OCT image are shown in Table 1. The deep learning model to predict BCVA at 12 months from the baseline OCT image had R 2 = 0.33, 0.75, and 0.58 and RMSE = 14.16, 11.27, and 13.25 for study eyes, fellow eyes, and all eyes, respectively (Table 7; FIG. 6A , FIG. 6B , FIG. 6C ). The deep learning model to predict BCVA at 12 months from the baseline OCT image had MD = -1.63 letters, 95% LOA, -29.48 to 26.22 letters for study eyes, MD = -2.31 letters, 95% LOA, -29.96 to 25.33 letters for fellow eyes, and MD = -1.97 letters, 95% LOA, -29.67 to 25.73 letters for all eyes FIG. 4A , FIG. 4B , FIG. 4C ). The multivariate linear model predicting 12-month BCVA from baseline OCT image and 12-month BCVA from baseline BCVA both predicted 12-month BCVA with R 2 = 0.40, 0.88, and 0.68 for study eyes, fellow eyes, and all eyes, respectively (Table 7).

[0173] Classification models to predict BCVA < 69 letters (Snellen equivalent, 20 / 40). The deep learning model to predict BCVA < 69 letters at month 12 from baseline OCT images had AUC = 0.80, 0.92, and 0.87 for the study eye, the fellow eye, and all eyes, respectively (Table 8; FIG. 7A , FIG. 7B , FIG. 7C ). In the study eye, the deep learning model to predict BCVA < 69 letters at month 12 from baseline OCT images had AUPRC = 0.74 with a class balance of 58 positive eyes and 68 negative eyes. In the fellow eye, the deep learning model to predict BCVA < 69 letters at month 12 from baseline OCT images had AUPRC = 0.97 with a class balance of 88 positive eyes and 37 negative eyes. In all eyes, the deep learning model to predict BCVA < 69 letters at month 12 from baseline OCT images had AUPRC = 0.91 with a class balance of 146 positive eyes and 105 negative eyes.

[0174] Classification models to predict BCVA < 59 letters (Snellen equivalent, 20 / 60). The deep learning model to predict BCVA < 59 letters at month 12 from baseline OCT images had AUC = 0.84, 0.93, and 0.89 for the study eye, the fellow eye, and all eyes, respectively (Table 8). In the study eye, the deep learning model to predict BCVA < 59 letters at month 12 from baseline OCT images had AUPRC = 0.90 with a class balance of 83 positive eyes and 43 negative eyes. In the fellow eye, the deep learning model to predict BCVA < 59 letters at month 12 from baseline OCT images had AUPRC = 0.98 with a class balance of 101 positive eyes and 24 negative eyes. In all eyes, the deep learning model to predict BCVA < 59 letters at month 12 from baseline OCT images had AUPRC = 0.95 with a class balance of 184 positive eyes and 67 negative eyes.

[0175]

[0176] BCVA, best-corrected visual acuity, OCT image, optical coherence tomography; CI, 95% confidence interval; *, P < 0.001 for the null hypothesis compared to the univariate model; **, P < 0.001 for each coefficient in the multivariate model.

[0177] Table 7. Performance of linear models for regression of BCVA at month 12 from baseline OCT image and baseline letters

[0178]

[0179] AUC, area under the receiver operating characteristic curve; BCVA, best-corrected visual acuity; OCT image, optical coherence tomography; Cl, 95% confidence interval.

[0180] Table 8. Performance of deep learning models for binary classification of BCVA < 69 letters (Snellen equivalent, 20 / 40), < 59 letters (Snellen equivalent, 20 / 60), and < 38 letters (Snellen equivalent, 20 / 200) at 12 months from baseline OCT image.

[0181] Classification model predicting BCVA < 38 letters (Snellen equivalent, 20 / 200). The deep learning model to predict BCVA < 38 letters at 12 months from baseline OCT image had an AUC of 0.77, 0.96, and 0.89 for the study eye, the fellow eye, and all eyes, respectively (Table 8). In the study eye, the deep learning model to predict BCVA < 38 letters at 12 months from baseline OCT image had an AUPRC = 0.97 with a class balance of 114 positive eyes and 12 negative eyes. In the fellow eye, the deep learning model to predict BCVA < 38 letters at 12 months from baseline OCT image had an AUPRC = 0.99 with a class balance of 109 positive eyes and 16 negative eyes. In all eyes, the deep learning model to predict BCVA < 38 letters at 12 months from baseline OCT image had an AUPRC = 0.98 with a class balance of 223 positive eyes and 28 negative eyes.

[0182] VI. A. 3. Discussion

[0183] As shown by the results above, deep learning models were able to predict BCVA from OCT images of subjects with neovascular AMD. The predictive accuracy of the derived models was highest in the fellow eye, with a correlation of about 0.92 (R2= 0.84, RMSE = 9.01; Table 3) between the mean predicted BCVA and the mean observed BCVA. In the study eye, moderate to strong correlations were also seen in the BCVA outcomes, with a correlation of about 0.49 (R2= 0.24, RMSE = 11.55; Table 3) at baseline (pre-treatment) and about 0.79 (R2= 0.62, RMSE = 10.54; Table 3) at 24 months (post-treatment). 2 2 2

[0184] ​​​To benchmark the presented model results (app.RMSE 10 letter errors), predictions were generated based on a simple linear regression of BCVA at baseline to BCVA in the HARBOR study using the baseline BCVA. The mean number of days between the screening and baseline visits was 8.3 days with a mean change of 0.6 letters; and, for this BCVA to BCVA prediction, the RMSE (from the regression) was 6.1 letters. The average error of 6.1 letters can represent an information limit inherent to the BCVA observations in the HARBOR study. Prior work on the intermittent repeatability of visual acuity scores in neovascular AMD reported an error of approximately 12 letters. Notably, the mean BCVA improvement with VEGF treatment in the HARBOR trial was reported to be 7.6 to 9.1 letters. The quality of the deep learning prediction model used in this example (which makes predictions from OCT images to BCVA) should be compared to the aforementioned (information) limit.

[0185] The results demonstrate that a mapping (defined by the deep learning model) between retinal structure and visual function exists in neovascular AMD. As a result, OCT images can provide a means to indirectly measure visual function in the context of clinical research as well as evolving clinical practice settings such as telemedicine or home monitoring.

[0186] Due to the limited range of BCVA at baseline for the study eye, differences in model performance between the study eye and the fellow eye visual function predictions were expected. Specifically, the HARBOR eligibility criteria required the study eye to have some vision loss and subfoveal CNV at baseline with BCVA between 20 / 40 and 20 / 320 (Snellen equivalent). These criteria were not required for the fellow eye. Thus, the limited range of BCVA at baseline for the study eye (SD = 13.2; Table 1) reduced the dynamic range and resulted in a more challenging regression task compared to the task of predicting BCVA for the fellow eye, which had greater variability in BCVA at baseline (baseline SD = 22.9; Table 1). This observation is supported by the fact that the prediction accuracy of simultaneously predicting BCVA from OCT images increased in the study eye over the course of the trial while the variability of BCVA increased (Tables 1-3, 9). This increase in range post-treatment is consistent with effective treatment producing visual improvement in many subjects. If, by simulation, the variance of BCVA in the fellow eye was (artificially) restricted to the variance of BCVA in the study eye at baseline, i.e., restricted to SD = 13.2, then R 2 decreased from 0.80 to 0.33 FIG. 8 ). Similarly, if log(var(BCVA)) versus log(l-R2) is plotted for the study eye and the fellow eye at baseline, 6 months, 12 months, 18 months, and 24 months, the R2values are 0.80, 0.75, 0.73, 0.71, and 0.67, respectively.2 The resulting (best-fit) pattern is linear (with a slope of -1.22) and correlates with both regression and correlation models around R0. 2 Consistent with the theory ( FIG. 9 Furthermore, the estimated residual errors (RMSE) show that these models appear to have similar performance at each time point (Tables 3 and 9).

[0187]

[0188] BCVA, best corrected visual acuity; OCT image, optical coherence tomography; CI, 95% confidence interval; RMSE, root mean square error.

[0189] Table 9. Cross-validation results. Performance of the deep learning model used for BCVA regression on OCT images.

[0190] Determining the precise relationship between specific and measurable anatomical changes and visual acuity is challenging. The conventional approach to studying this relationship is to select one or more anatomical features and then analyze them to determine if any association with vision can be quantified. Therefore, this approach is limited by the ability of researchers to pre-determine a large set of potential retinal structures and features most likely to have a meaningful relationship with visual function. In contrast, deep learning-based algorithms (and particularly CNNs) do not require the identification of any anatomical features before quantitative analysis. Instead, deep learning algorithms evaluate OCT images as a whole and learn directly from the images to identify features that can most accurately predict outcomes of interest. Compared to the cross-validation results of a previously reported subset of 614 participants from the HARBOR trial (which reported that when predicting BCVA at month 12 using known imaging features and baseline BCVA), R... 2 =0.34), the regression deep learning model studied in this example achieved R on 924 subjects in the adjusted set. 2 =0.45, and R was achieved on 126 subjects in the internal validation test set. 2 =0.40. In principle, this could lead to the identification of previously overlooked anatomical features or combinations of features that are crucial to visual function.

[0191] A standalone deep learning model was able to predict the BCVA values ​​of the study eyes 12 months from the time of baseline OCT image measurement, with an isocorrelation of approximately 0.57 (R²). 2= 0.33; Table 7). Interestingly, it was noted that when added to a regression model already containing baseline BCVA (P < 0.001), the OCT image-based prediction (from baseline) remained highly statistically significant (P < 0.001). In this multivariate model, the two predictors provided roughly equal information about future visual function, with model R 2 = 0.40 (Table 7). If used as a stratification factor at baseline, this prediction model could translate into a smaller / shorter trial with the same statistical power. Initially, separate models were used for study eyes and fellow eyes, but surprisingly, there were no significant differences in predictive performance between the models trained in the study and the fellow eyes combined.

[0192] The use of deep learning models to predict BCVA can have meaningful clinical utility. The measurement of BCVA is often cumbersome, requiring specialized resources for accurate refraction testing. In fact, the ability to augment visual function measurements through computer vision-based OCT image analysis in retinal health assessments outside of the clinical setting would likely be valuable for screening and monitoring subjects. For example, it could aid in telemedicine for remote consultations, where physicians could use deep learning data about a subject’s current and future visual potential to support their clinical decisions. Furthermore, in clinical research, deep learning models that help predict future BCVA response could be used to support trial enrollment or trial stratification by focusing on individuals likely to benefit from treatment.

[0193] VI. B. Machine learning model for processing color fundus images to predict vision metrics

[0194] This example assessed whether deep learning could automatically predict BCVA from color fundus photograph (CFP) images from subjects with neovascular AMD. Specifically, a first deep learning regression model (which included a deep convolutional neural network and a linear activation function) was used to predict accurate BCVA from CFP images at chart distances of 2 meters (m) and 4 m. In addition, a second deep learning classification model (which included a deep convolutional neural network and a softmax activation function) was also used.

[0195] VI. B. 1. Methods and data

[0196] VI. B. 1. a. Data set sources

[0197] Prospectively collected BCVA measurements and CFP images from 707 subjects from the Phase 3 MARINA clinical study (NCT00056836) and 413 subjects from the Phase 3 ANCHOR clinical study (NCT00061594) were used. MARINA and ANCHOR complied with the principles of the Declaration of Helsinki and were in accordance with the Health Insurance Portability and Accountability Act. The protocol was approved by each institutional review board prior to study initiation and written informed consent was provided for all subjects regarding future medical research and analyses based on trial results.

[0198] In MARINA, 720 adult subjects with subfoveal choroidal neovascularization (CNV) secondary to neovascular AMD were enrolled provided they had BCVA in the study eye between 20 / 40 and 20 / 320 (Snellen equivalent) as measured using standard ETDRS charts and protocols. In ANCHOR, 426 adult subjects with subfoveal choroidal neovascularization (CNV) secondary to neovascular AMD were enrolled provided they had BCVA in the study eye between 20 / 40 and 20 / 320 (Snellen equivalent).

[0199] VI. B. 1. b. CFP images

[0200] A total of 36,541 images from MARINA and 33,591 images from ANCHOR were analyzed. The inner capture fields of F1M, F2, and F3M from left and right stereo views were included in the analysis, while the outer views of the eye (FR capture fields) were excluded. (It is worth noting that F4, F5, F6, and / or F7 fields can alternatively be used to predict visual acuity metrics). To remove extraneous information, the images were cropped to fit the circle produced by the camera lens. The image size was then re-set to 299x299x3 pixels. CFP images from MARINA were split into five folds at the subject level for model development via cross-validation. For both regression and classification tasks, the subjects in each fold remained constant.

[0201] VI. B. 1. c. Outcome variables for deep learning modeling

[0202] The BCVA outcomes of interest were: (1) BCVA in ETDRS letters at each visit, and (2) whether the specific BCVA value was <69 letters (Snellen equivalent of 20 / 40). The Snellen equivalent of 20 / 40 was chosen because it was considered to reflect a functionally meaningful level of visual impairment. The mean BCVA (± standard deviation [SD]) for eyes in the ANCHOR external validation test set was 55.7 ± 24.7 letters for subjects for whom BCVA was measured at 2 m distance, and 55.0 ± 25.2 letters for subjects for whom BCVA was measured at 4 m distance (Table 10).

[0203]

[0204] Table 10. Characteristics of the MARINA and ANCHOR datasets used for model development and testing, respectively.

[0205] VI. B. 1. d. Deep learning algorithm

[0206] The ability of the deep learning model to predict the following was evaluated: (1) accurate BCVA values in ETDRS letters from CFP images obtained at the same visit; and (2) BCVA <69 letters (Snellen equivalent, 20 / 40) from CFP images obtained at the same visit.

[0207] Deep learning modeling was performed using TensorFlow (1.14.0) and Keras (2.2.5) on an Nvidia V100 GPU and an Inception-ResNet-v2 CNN architecture that was trained to predict BCVA at the same visit as the CFP.

[0208] For the regression model, a global average pooling layer, a dropout layer (0.5), a dense layer (256), and a dense layer (1) were added to the base CNN model. To account for the distance at which BCVA was measured, a corresponding figure distance of 2 m or 4 m was concatenated to the final dense layer ( FIG. 10 ). The loss function was mean squared error. For each of the five-fold cross-validation, the model was initialized with pre-trained weights on the ImageNet dataset and trained for 2 epochs with the Adam optimizer with non-trainable base model layers, and an additional 200 epochs with the RAdam optimizer with trainable base model layers.

[0209] For the classification model, the architecture remained the same as the regression model except that the last layer was a dense layer with a softmax activation function using a sparse categorical cross-entropy loss function (2). The model was initialized with the weights from the regression model from each fold. The model was trained for 3 epochs using the Adam optimizer with the non-trainable base model layers.

[0210] VI. B. 1. e. Evaluation of deep learning model

[0211] The model weights from the epoch with the lowest validation loss were selected from each cross-validation fold. The metric used to evaluate the model fit was calculated as the mean of the predictions generated for the eye at each visit across each of the five-fold cross-validation on the ANCHOR out-of-sample external validation test set (Table 11). In addition, results from the MARINA five-fold cross-validation adjustment set were reported (Table 11). 2 The value was used to benchmark the deep learning regression model, while the area under the receiver operating characteristic curve (AUC) was used to assess the classification performance of the deep learning model.

[0212] VI. B. 2. Results

[0213] Regression model results. The regression model to predict BCVA at chart distance of 2 m had R2= 0.56 (95% CI: 0.54, 0.57) in the MARINA development set and R2= 0.59 (95% CI: 0.57, 0.60) in the ANCHOR external validation test set (Table 11, 2 FIG. 11A ). The regression model to predict BCVA at chart distance of 4 m had R2= 0.57 (95% CI: 0.55, 0.60) in the MARINA development set and R2= 0.60 (95% CI: 0.57, 0.63) in the ANCHOR external validation test set (Table 11, 2 2 FIG. 11B FIG. 11A and FIG. 11B shows the performance of the deep learning regression model that analyzed color fundus photograph images to predict BCVA in the ANCHOR external validation test set. FIG. 11A shows the actual BCVA versus the predicted BCVA at chart distance of 2 meters. R 2 = 0.59. FIG. 11B shows the actual BCVA versus the predicted BCVA at chart distance of 4 meters. R 2 = 0.60.

[0214] ​​​​Classification model results. The classification model to predict BCVA <69 letters (Snellen equivalent 20 / 40) at chart distance 2 m had AUC = 0.86 (95% CI: 0.85, 0.87) in the MARINA development test set and AUC = 0.86 (95% CI: 0.85, 0.87) in the ANCHOR external validation test set (Table 11, FIG. 12A ). The classification model to predict BCVA <69 letters (Snellen equivalent 20 / 40) at chart distance 4 m had AUC = 0.87 (95% CI: 0.85, 0.88) in the MARINA development test set and AUC = 0.88 (95% CI: 0.86, 0.90) in the ANCHOR external validation test set (Table 11, FIG. 12B ). FIG. 12A and FIG. 12B shows the performance of a deep learning classification model that analyzes color fundus photograph images to predict BCVA <69 letters (Snellen equivalent <20 / 40) in the ANCHOR external validation test set. FIG. 12A shows the area under the receiver operating characteristic curve (AUC) at chart distance 2 meters. AUC = 0.86. Figure 12b shows the AUC at chart distance 4 meters. AUC = 0.88.

[0215] BCVA@2m BCVA@4m MARINA BCVA R 2 ]]> 0.56 (95% CI: 0.54, 0.57) 0.57 (95% CI: 0.55, 0.60) ANCHOR BCVA R 2 ]]> 0.59 (95% CI: 0.57, 0.60) 0.60 (95% CI: 0.57, 0.63) MARIA BCVA < 69 AUC 0.86 (95% CI: 0.85, 0.87) 0.87 (95% CI: 0.85, 0.88) ANCHOR BCVA < 69 AUC 0.86 (95% CI: 0.85, 0.87) 0.88 (95% CI: 0.86, 0.90)

[0216] Table 11. Model performance on the adjustment (MARINA) and test (ANCHOR) datasets.

[0217] VI. B. 3) Discussion of described methods.

[0218] The results demonstrate that a neural network can learn quantitative relationships between retinal structure and visual function in subjects with neovascular age-related macular degeneration.

[0219] VII. Computing system

[0220] FIG. 13A network 1300 of computing systems is shown, which can be configured to perform part or all of one or more actions and / or one or more methods disclosed herein. The network 1300 can include one or more eye imaging systems 1305 configured to collect one or more images of a subject’s eye. The eye imaging system can include one or more techniques disclosed herein (e.g., optical coherence tomography as disclosed in Section II.A.1, or color fundus photography as disclosed in Section II.A.2). The eye imaging system can include optical components configured to, for example, collect OCT images or color fundus photographs. For example, the eye imaging system can include an interferometer (e.g., a Michelson-type interferometer), a light source (e.g., low coherence, wide bandwidth), and a beamsplitter. As another example, the eye imaging system can include a fundus camera. The eye imaging system can include a computing system (e.g., having one or more processors, one or more memories, and / or one or more transmitters) to store and / or transmit the images (or processed versions thereof).

[0221] The network 1300 can include one or more visual function assessment systems 1310, which can include one or more computing systems (e.g., having one or more processors, one or more memories, or one or more transmitters) configured to collect and transmit visual acuity metrics generated based on a subject’s responses to observing visual stimuli, such as an eye chart or eye test card. The visual acuity metrics can include the metrics disclosed in Section II.C.1. The visual acuity metrics can be determined based on the techniques disclosed herein. In some cases, the visual function assessment system presents the visual stimuli on a screen of the system. In some cases, the visual stimuli are presented separately (e.g., via a physical chart or card).

[0222] The machine learning model system 1315 can include one or more computing systems (e.g., having one or more processors, one or more memories, one or more transmitters, and / or one or more receivers). The machine learning model system 1315 can be, in part or in its entirety, a cloud computing system. The machine learning model system 1315 can include one or more servers. In some cases, the machine learning model system 1315 is or is included within a subject device and / or a subject’s medical device (e.g., a wearable device, a smart phone, etc.).

[0223] The machine learning model system 1315 can be configured to train and / or use a machine learning model, which can include one or more pre-processing functions (e.g., as disclosed in Section II.B.) and one or more neural networks (e.g., having the architectures and / or properties disclosed in Section III.A.).

[0224] The machine learning model system 1315 can train a machine learning model using, for example, the techniques disclosed in Section III.B. The machine learning model system 1315 can receive training data (e.g., disclosed in Section II.C) for each training subject in a set of training subjects. For each training subject, the training data can include one or more images of one or both eyes (from the eye imaging system 1305) and a vision score or metric corresponding to the one or both eyes (from the visual function assessment system 1310). In some cases, for each of the one or both eyes, multiple vision metrics corresponding to different time points (e.g., relative to the time at which the eye images were collected) are received. For example, one metric can correspond to a visual function test (e.g., reading of an eye chart) performed on the same day as the eye images were collected, while another metric corresponds to a visual function test performed approximately 6 months or approximately 1 year after the eye images were collected. The machine learning model system 1315 can use the images and vision metrics to train a machine learning model (e.g., using a deep convolutional network) to predict a vision metric (or another type of vision metric, such as a binary indicator) based on images of the eye. In some cases, the model includes one or more pre-processing functions to pre-process the images. Training the model can include learning a set of parameters.

[0225] The machine learning model system 1315 can then receive another image corresponding to another subject (e.g., from another eye imaging system 1305) and can use the trained model to predict a current and / or future vision metric for the subject associated with the other image.

[0226] The machine learning model system 1315 can transmit the predicted vision metric to one or more user devices 1320 (e.g., along with an identifier for the subject). The one or more user devices 1320 can correspond to, for example, a care provider for the subject or the subject him / herself. The user of the one or more user devices 1320 can use the result in the manner disclosed in, for example, Section IV.

[0227] In some cases, two or more of the eye imaging system 1305, the visual function assessment system 1310, or the user device 1320 are owned by the same entity, co-located, and / or included within the same computing system (e.g., at an office of a care provider). In some cases, a device of the subject (e.g., a smartphone) includes at least a portion of each of the two or more depicted components. For example, an image of the subject’s eye can be collected using an accessory of the smartphone, the trained machine learning model can be run locally on the smartphone, and the predicted vision can be provided to the smartphone.

[0228] VIII. Exemplary embodiments

[0229] A first example includes a method for predicting vision based on image processing. The method includes accessing an image of at least a portion of an eye of a subject; inputting the image into a machine learning model to determine a predicted vision metric corresponding to a predicted vision, wherein the machine learning model includes: a set of parameters determined using: a set of training images, each training image in the set of training images depicting at least a portion of an eye of a training subject in a set of training subjects; and a set of labels identifying an observed vision of each training subject in the set of training subjects; and further using a function that relates the image and the parameters to a vision metric (e.g., wherein the function is learned to relate training images to observed vision metrics, and wherein the function can be used to relate non-training images to predicted vision metrics); and returning the predicted vision metric.

[0230] A second example includes the first example, wherein the image is an optical coherence tomography (OCT) image.

[0231] A third example includes the first example, wherein the image is a color fundus photograph image.

[0232] A fourth example includes any of the first through third examples, wherein the predicted vision metric corresponds to a predicted vision at a baseline date when the image was captured by an imaging device.

[0233] A fifth example includes any of the first through third examples, wherein the predicted vision metric includes one or more numbers representing a predicted vision at a date at least 6 months from a baseline date when the image was captured.

[0234] A sixth example includes any of the first through third examples, wherein the predicted vision metric is a binary value representing a prediction of whether the subject has worse than a threshold vision value at a baseline date when the image was captured by an imaging device.

[0235] A seventh example includes any of the first through third examples, wherein the predicted vision metric is a binary value representing a prediction of whether the subject has worse than a threshold vision value at a date at least 6 months from a baseline date when the image was captured by an imaging device.

[0236] The eighth example embodiment includes the sixth or seventh example embodiment, wherein the threshold value is equal to Snellen fraction 20 / 160, 20 / 80, or 20 / 40.

[0237] The ninth example embodiment includes any of the first through eighth example embodiments, further comprising, in response to the predicted visual acuity metric:

[0238] providing a recommendation that the subject receive a medication.

[0239] The tenth example embodiment includes any of the first through ninth example embodiments, wherein the image is a three-dimensional image, and wherein the method further comprises:

[0240] slicing the image into a plurality of two-dimensional slices, each of the slices being captured at a different scan center angle and the slices being offset from each other by a different number of pixels; and

[0241] wherein inputting the image into the model includes inputting the slices into the model.

[0242] The eleventh example embodiment includes any of the first through tenth example embodiments, wherein the model is a deep learning model.

[0243] The twelfth example embodiment includes any of the first through eleventh example embodiments, wherein the model is or includes a convolutional neural network.

[0244] The thirteenth example embodiment includes any of the first through twelfth example embodiments, wherein the model uses a set of convolutional kernels.

[0245] The fourteenth example embodiment includes any of the first through thirteenth example embodiments, wherein the model includes a ResNet model.

[0246] The fifteenth example embodiment includes any of the first through fourteenth example embodiments, wherein the model is an initial model.

[0247] The sixteenth example embodiment includes any of the first through fifteenth example embodiments, wherein the predicted visual acuity metric corresponds to a predicted visual acuity of an eye of the subject.

[0248] The seventeenth example embodiment includes any of the first through sixteenth example embodiments, wherein the predicted visual acuity metric corresponds to a predicted corrected visual acuity that predicts an eye of the subject’s vision when wearing eyeglasses or contact lenses.

[0249] The eighteenth example embodiment includes any of the first through seventeenth example embodiments, wherein the predicted vision metric corresponds to a predicted best-corrected visual acuity of the subject’s eye.

[0250] The nineteenth example embodiment includes any of the first through seventeenth example embodiments, wherein the subject was previously diagnosed with age-related macular degeneration at the time the images were collected.

[0251] The twentieth example embodiment includes any of the first through seventeenth example embodiments, wherein the subject was previously diagnosed with neovascular age-related macular degeneration at the time the images were collected.

[0252] The twenty-first example embodiment includes any of the first through seventeenth example embodiments, wherein the subject was previously diagnosed with atrophic age-related macular degeneration at the time the images were collected.

[0253] The twenty-second example embodiment includes any of the first through seventeenth example embodiments, wherein each of the training subjects was previously diagnosed with age-related macular degeneration prior to the training images in the set of images having been collected.

[0254] The twenty-third example embodiment includes any of the first through seventeenth example embodiments, wherein each of the training subjects was previously diagnosed with neovascular age-related macular degeneration prior to the training images in the set of images having been collected.

[0255] The twenty-fourth example embodiment includes any of the first through seventeenth example embodiments, wherein each of the training subjects was previously diagnosed with atrophic age-related macular degeneration prior to the training images in the set of images having been collected.

[0256] The twenty-fifth example embodiment includes any of the first through twenty-fourth example embodiments, further comprising:

[0257] training the model using the set of training images and the set of labels.

[0258] The twenty-sixth example embodiment includes any of the first through twenty-fifth example embodiments, wherein the machine learning model includes one or more pre-processing functions and one or more neural networks.

[0259] A twenty-seventh example includes any of the first through twenty-fifth examples, wherein the image of at least a portion of the eye includes a pre-processed version of the eye, the pre-processed version of the eye being generated by applying one or more pre-processing functions to an original image of at least a portion of the subject’s eye.

[0260] A twenty-eighth example includes any of the twenty-sixth or twenty-seventh examples, wherein the one or more pre-processing functions include a function that flattens the image.

[0261] A twenty-ninth example includes any of the twenty-sixth or twenty-seventh examples, wherein the one or more pre-processing functions include a function that generates one or more B-scan images or one or more C-scan images.

[0262] A thirtieth example includes any of the twenty-sixth or twenty-ninth examples, wherein the one or more pre-processing functions include a cropping function.

[0263] A thirty-first example includes any of the twenty-sixth or twenty-ninth examples, wherein the predicted vision metric corresponds to a predicted vision of the subject’s eye at a first distance, and wherein the method further comprises:

[0264] determining, using the machine learning model, another predicted vision metric corresponding to a predicted vision of a same or different eye of the subject; and

[0265] returning the other predicted vision metric.

[0266] A thirty-second example includes a method of stratifying a clinical study, the method comprising: performing, for each subject in a set of subjects, the method for predicting vision based on image processing according to any of the first through thirty-first examples; determining, for each subject in the set of subjects, whether the subject is eligible to participate in a clinical study, wherein the determination is based on whether each eligibility criterion in a set of eligibility criteria relating to the subject is satisfied, and wherein an assessment of a particular eligibility criterion in the set of eligibility criteria is performed using the predicted vision of the subject; and performing the clinical study with a subset of the set of subjects, wherein each subject in the subset is determined to be eligible to participate in the clinical study.

[0267] A thirty-third example includes the thirty-second example, further comprising: performing the clinical study according to an allocation of subjects to a first group of subjects and a second group of subjects.

[0268] A thirty-fourth example includes a method comprising: selecting a treatment among a set of potential treatments for the subject based on the predicted vision metric determined in accordance with the methods according to the first through thirty-first example; and outputting an identification of the treatment.

[0269] A thirty-fifth example includes: selecting a treatment among a set of potential treatments for the subject based on the predicted vision metric determined in accordance with the first through thirty-first example; and treating the subject with the treatment.

[0270] A thirty-sixth example includes the thirty-fourth or thirty-fifth example, wherein each of at least some of the training subjects has received the treatment prior to participating in a vision visual test from which the label corresponding to the training subject was determined.

[0271] A thirty-seventh example includes the thirty-fourth or thirty-fifth example, wherein each of the training subjects has received the treatment prior to participating in a vision visual test from which the label corresponding to the training subject was determined.

[0272] A thirty-eighth example includes the thirty-fourth or thirty-fifth example, wherein each of at least some of the training subjects has received another treatment different from the treatment prior to participating in a vision visual test from which the label corresponding to the training subject was determined, wherein the subject was previously receiving the other treatment.

[0273] A thirty-ninth example includes the thirty-fourth or thirty-fifth example, wherein each of the training subjects has received another treatment different from the treatment prior to participating in a vision visual test from which the label corresponding to the training subject was determined, wherein the subject was previously receiving the other treatment.

[0274] A fortieth example includes: one or more data processors; and a non-transitory computer-readable storage medium containing instructions that, when executed by the one or more data processors, cause the one or more data processors to perform part or all of one or more methods disclosed herein.

[0275] A forty-first example includes a computer program product tangibly embodied in a non-transitory machine-readable storage medium including instructions configured to cause one or more data processors to perform part or all of one or more methods disclosed herein.

[0276] IX. Additional considerations

[0277] Some of the disclosure herein relates to a subject (e.g., identifying vision of a subject, predicting vision of a subject, treating a subject, etc.). It will be appreciated that, in some embodiments, some or all of the disclosure relates to a particular eye of a subject.

[0278] Some embodiments of the disclosure include a system comprising one or more data processors. In some embodiments, the system includes a non-transitory computer- readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform part or all of one or more methods disclosed herein and / or part or all of one or more processes. Some embodiments of the disclosure include a computer program product tangibly embodied in a non-transitory machine- readable storage medium including instructions configured to cause one or more data processors to perform part or all of one or more methods disclosed herein and / or part or all of one or more processes.

[0279] The terminology and expressions used herein are used as descriptive terminology and not as limiting terminology unless expressly stated otherwise. In using such terminology and expressions, no limitation is intended to be implied, unless otherwise explicitly limited. It will be appreciated that any modifications and variations of the concepts herein disclosed can be resorted to by those skilled in the art, and that such modifications and variations are considered to be within the scope of the application as defined by the appended claims.

[0280] This description provides preferred exemplary embodiments only and is not intended to limit the scope, applicability or configuration of the disclosure. Rather, the subsequent description of preferred exemplary embodiments will provide those skilled in the art with an enabling description for implementing various embodiments. It should be understood that various changes can be made in the function and arrangement of elements without departing from the spirit and scope as expressed in the appended claims.

[0281] In the following description, specific details are set forth to provide a thorough understanding of the embodiments. However, it will be appreciated that the embodiments can be practiced without these specific details. For example, circuits, systems, networks, processes and other components can be shown as components in block diagram form, rather than in detail in order to avoid obscuring embodiments. In other instances, well-known circuits, processes, algorithms, structures and techniques have not been shown in detail in order to avoid obscuring embodiments.

Claims

1. A method for predicting vision based on image processing, the method comprising: accessing an image of at least a portion of an eye of a subject, wherein the image is captured by an imaging device at a first date; inputting the image into a machine learning model to determine: a first predicted vision metric that predicts vision of the subject, the first predicted vision metric being determined based on a vision assessment performed on the subject at the first date at which the image is captured; and a second predicted vision metric that predicts vision of the subject, the second predicted vision metric being determined based on a vision assessment performed on the subject at a second date at a particular interval after the first date, wherein the machine learning model comprises: a set of parameters determined using: a set of training images corresponding to a set of training subjects, wherein each training image in the set of training images depicts at least a portion of an eye of a corresponding training subject in the set of training subjects; and a set of labels that identifies an observed vision metric for each training subject in the set of training subjects; wherein the observed vision metric comprises a numerical vision value and / or an identification of a numerical vision range; and wherein the observed vision metric is associated with the subject at a time at which the image is collected or at a time after the image is collected; a function that associates the image and the set of parameters with the first predicted vision metric and the second predicted vision metric; and returning the first predicted vision metric and the second predicted vision metric.

2. The method of claim 1, wherein the image is an optical coherence tomography (OCT) image.

3. The method of claim 1, wherein the image is a color fundus photograph image.

4. The method of any one of claims 1-3, wherein the second date is 12 months after the first date.

5. The method of any one of claims 1 to 3, wherein the second predicted measure of visual acuity comprises one or more numbers representing a predicted visual acuity at the second date, wherein, the second date is at least 6 months after the first date at which the image is captured.

6. The method of any one of claims 1-3, wherein the first predicted vision metric is a binary value that predicts whether the subject’s vision at the first date is worse than a threshold vision value.

7. The method of any one of claims 1 to 3, wherein the second predicted measure of visual acuity is a binary value that predicts whether the subject has worse visual acuity than a threshold visual acuity value at the second date, wherein, the second date is at least 6 months after the first date.

8. The method of claim 6 or 7, wherein the threshold vision value is equal to a Snellen fraction of 20 / 160, 20 / 80, or 20 / 40.

9. The method of any one of claims 1-8, further comprising, in response to the first predicted vision metric and the second predicted vision metric: providing a recommendation that the subject receive a medication.

10. The method of any one of claims 1-9, wherein the image is a three-dimensional image, and wherein the method further comprises: slices each of which is captured at a different scan center angle and which are offset from each other by different numbers of pixels; and wherein inputting the image into the model comprises inputting the slices into the model.

11. The method of any one of claims 1 to 10, wherein the model is a deep learning model.

12. The method of any one of claims 1 to 11, wherein the model is or comprises a convolutional neural network.

13. The method of any one of claims 1 to 12, wherein the model uses a set of convolutional kernels.

14. The method of any one of claims 1 to 13, wherein the model comprises a ResNet model.

15. The method of any one of claims 1 to 14, wherein the model is an Inception model.

16. The method of any one of claims 1 to 15, wherein the first predicted vision metric and the second predicted vision metric correspond to a predicted visual acuity of the subject’s eye.

17. The method of any one of claims 1 to 16, wherein the first predicted vision metric and the second predicted vision metric correspond to a predicted corrected visual acuity that predicts the visual acuity of the eye subject when wearing eyeglasses or contact lenses.

18. The method of any one of claims 1 to 17, wherein the first predicted vision metric and the second predicted vision metric correspond to a predicted best-corrected visual acuity of the subject’s eye.

19. The method of any one of claims 1 to 17, wherein the subject was previously diagnosed with age-related macular degeneration at the time the images were collected.

20. The method of any one of claims 1 to 17, wherein the subject was previously diagnosed with neovascular age-related macular degeneration at the time the images were collected.

21. The method of any one of claims 1 to 17, wherein the subject was previously diagnosed with atrophic age-related macular degeneration at the time the images were collected.

22. The method of any one of claims 1 to 17, wherein each of the training subjects was previously diagnosed with age-related macular degeneration prior to the training images in the set of images having been collected.

23. The method of any one of claims 1 to 17, wherein each of the training subjects was previously diagnosed with neovascular age-related macular degeneration prior to the training images in the set of images having been collected.

24. The method of any one of claims 1 to 17, wherein each of the training subjects was previously diagnosed with atrophic age-related macular degeneration prior to the training images in the set of images having been collected.

25. The method of any one of claims 1-24, further comprising: training the model using the set of training images and the set of labels.

26. The method of any one of claims 1-25, wherein the machine learning model comprises one or more pre-processing functions and one or more neural networks.

27. The method of any one of claims 1-25, wherein the image of at least a portion of the eye comprises a pre-processed version of the eye, the pre-processed version of the eye being generated by applying one or more pre-processing functions to an original image of at least a portion of the subject’s eye.

28. The method of claim 26 or claim 27, wherein the one or more pre-processing functions comprise a function that flattens an image.

29. The method of any one of claims 26-28, wherein the one or more pre-processing functions comprise a function that generates one or more B-scan images or one or more C-scan images.

30. The method of any one of claims 26-29, wherein the one or more pre-processing functions comprise a cropping function.

31. The method of any one of claims 26-29, wherein the first predicted vision metric corresponds to a predicted vision of the subject’s eye at a first distance, and wherein the method further comprises: using the machine learning model to determine another predicted vision metric corresponding to a predicted vision of the same or a different eye of the subject; and returning the other predicted vision metric.

32. A method of stratifying a clinical study, the method comprising: performing, for each subject in a set of subjects, the method for predicting vision based on image processing of any one of claims 1-31; determining, for each subject in the set of subjects, whether the subject is eligible to participate in a clinical study, wherein the determination is based on whether each eligibility criterion in a set of eligibility criteria pertaining to the subject is satisfied, and wherein an assessment of a particular eligibility criterion in the set of eligibility criteria is performed using the predicted vision of the subject; and performing the clinical study with a subset of the set of subjects, wherein each subject in the subset is determined to be eligible to participate in the clinical study.

33. The method of claim 32, further comprising: performing the clinical study according to an assignment of subjects to a first group of subjects and a second group of subjects.

34. A system comprising: one or more data processors; and a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform the method of any one of claims 1-33. ​ 35. A computer program product tangibly embodied in a non-transitory machine- readable storage medium including instructions configured to cause one or more data processors to perform a method according to any one of claims 1-33.