Using deep learning to process eye images to predict visual acuity

Deep learning models analyze OCT and fundus images to predict visual acuity, addressing the limitations of traditional methods by identifying specific anatomical-visual relationships, enhancing the precision of AMD progression prediction and treatment planning.

JP7750827B2Active Publication Date: 2025-10-07GENENTECH INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2022506613
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-03-16
Filing Date
2020-07-31
Publication Date
2025-10-07
Estimated Expiration
2040-07-31

AI Technical Summary

Technical Problem

Current technologies lack the ability to reliably predict the severity and progression of retinal diseases like neovascular age-related macular degeneration (AMD) based on ocular anatomy, as traditional correlation analyses are limited by the need to pre-specify features for analysis and often result in aggregated measures that do not have a specific relationship to visual outcomes.

Method used

A method using machine learning, specifically deep learning models such as convolutional neural networks, to analyze optical coherence tomography (OCT) and color fundus images to predict visual acuity by identifying novel relationships between anatomical and visual parameters, enabling precise prediction of current and future visual acuity.

Benefits of technology

The method provides a more reliable prediction of visual acuity, facilitating treatment decisions and clinical trial stratification by accurately forecasting visual function based on eye anatomy, thereby improving monitoring and treatment strategies for retinal diseases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007750827000012
    Figure 0007750827000012
  • Figure 0007750827000013
    Figure 0007750827000013
  • Figure 0007750827000014
    Figure 0007750827000014
Patent Text Reader

Abstract

Systems and methods disclosed herein relate to using machine learning models to process ocular inputs from a subject and predict the subject's current or future visual acuity. The subject may have been diagnosed with age-related macular degeneration. The predicted current or future visual acuity can be used (for example) to facilitate diagnosis of the subject (e.g., with a particular type of age-related macular degeneration), facilitate identification of a treatment strategy for the subject, and / or facilitate design of a clinical trial.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application claims the benefit of and priority to the following U.S. provisional patent applications: 62 / 882,354, filed August 2, 2019, 62 / 907,014, filed September 27, 2019, 62 / 940,989, filed November 27, 2019, and 62 / 990,354, filed March 16, 2020. Each of these applications is incorporated herein by reference in its entirety for all purposes. [Background technology]

[0002] Retinal diseases, such as neovascular age-related macular degeneration (NEOAMD), are characterized by pathophysiological and anatomical changes that can interfere with vision and lead to permanent vision loss. Age-related macular degeneration (AMD) is an eye disease that affects a person's vision. There are two types of AMD: dry and wet. Dry AMD is the more common and milder form of AMD. It usually progresses gradually and slowly affects vision. Wet AMD, also known as neovascular age-related macular degeneration, is a more advanced version of AMD. One risk factor for neovascular AMD is age, which typically occurs in people over the age of 50. Smoking also increases a person's likelihood of developing neovascular AMD by two to five times. Other risk factors for neovascular AMD include obesity, genetics, race, and gender. People with a body mass index (BMI) over 30 are 2.5 times more likely to develop neovascular AMD than those with a lower BMI. Additionally, a family history of neovascular AMD increases the likelihood of developing neovascular AMD. Women, Caucasians, and people with light-colored eyes are also at higher risk of developing neovascular AMD. While risk factors can provide information about the likelihood of developing neovascular AMD, no current technology can reliably predict the severity and rate of degeneration of the disease.

[0003] Therefore, caregivers often rely on frequent monitoring to determine whether to initiate and / or modify a given treatment (e.g., intravitreal anti-vascular endothelial growth factor, which results in improved visual acuity in some subjects). One approach to objectively assessing visual acuity is to determine which letters on an eye chart (e.g., a Snellen eye chart or LogMAR) are correctly identified by the individual. The eye chart can include letters of different sizes (e.g., different sizes are presented on different lines on the eye chart), and visual acuity can be determined based on determining which size letters are correctly identified by the observer. When the observer views the eye chart using glasses or contact lenses, the visual acuity metric is typically referred to as "best-corrected visual acuity."

[0004] One reason that monitoring corrected or best-corrected visual acuity can be beneficial is that various visual diseases and / or conditions can cause changes in the eye that cannot be corrected with glasses or contact lenses. For example, excess fluid (e.g., in the macula) can cause blurred vision that cannot be corrected with glasses or contact lenses. On the other hand, many other types of vision loss that occur naturally as a result of aging can be corrected with glasses or contact lenses. Therefore, corrected or best-corrected visual acuity can be an indicator of the presence, progression, and / or stage of a disease. Summary of the Invention [Means for solving the problem]

[0005] Various ocular diseases (e.g., including AMD) cause anatomical changes and further decline in visual function, yet precise relationships predicting current or future visual acuity based on ocular anatomy have yet to be identified. Traditional correlation analyses are limited in their ability to detect novel relationships between anatomical and visual parameters by the need to identify and pre-specify candidate sets of features for analysis, a process limited by the insight of human researchers. Furthermore, traditional methods that pre-specify features such as central segmental thickness (CST) and central foveal thickness (CFT) typically result in aggregated measures of retinal health that may not have a sufficiently specific relationship to visual outcome. For example, in the HARBOR trial data, features derived from intraretinal fluid and total retinal thickness were associated with R 2 Correlated with baseline BCVA at =0.21.

[0006] Therefore, there is a desire to more reliably predict visual function based on anatomy.

[0007] The present disclosure is described in conjunction with the accompanying drawings, in which: In an embodiment of the present invention, for example, the following items are provided: (Item 1) A method for predicting visual acuity based on image processing, comprising: accessing an image of at least a portion of the subject's eye; inputting the image into a machine learning model to determine a predicted visual acuity metric corresponding to a predicted visual acuity, wherein the machine learning model: A set of parameters, a set of training images, each of the set of training images depicting at least a portion of an eye of a training subject of a set of training subjects; a set of labels identifying the observed visual acuity for each of the set of training subjects; and a set of parameters determined using a function relating the image and the parameters to a visual acuity metric; returning the predicted visual acuity metric; A method comprising: (Item 2) Item 10. The method of item 1, wherein the image is an optical coherence tomography (OCT) image. (Item 3) Item 10. The method according to item 1, wherein the image is a color fundus photograph image. (Item 4) 4. The method of any one of items 1 to 3, wherein the predicted visual acuity metric corresponds to the predicted visual acuity at a baseline date when the image was captured by an imaging device. (Item 5) 4. The method of any one of items 1 to 3, wherein the predicted visual acuity metric comprises one or more numbers, the one or more numbers representing predicted visual acuity for a date at least six months from the baseline date at which the image was captured. (Item 6) 4. The method of any one of items 1 to 3, wherein the predicted visual acuity metric is a binary value representing a prediction as to whether the subject's visual acuity on a baseline date when the image was captured by an imaging device is worse than a threshold visual acuity value. (Item 7) 4. The method of any one of items 1 to 3, wherein the predicted visual acuity metric is a binary value representing a prediction as to whether the subject's visual acuity at least 6 months from a baseline date when the image was captured by an imaging device will be worse than a threshold visual acuity value. (Item 8) 8. The method according to item 6 or 7, wherein the threshold value corresponds to a Snellen fraction of 20 / 160, 20 / 80 or 20 / 40. (Item 9) in response to the predicted visual acuity metric; 9. The method of any one of items 1 to 8, further comprising providing a recommendation that the subject undergo pharmacological treatment. (Item 10) the image is a three-dimensional image, and the method comprises: further comprising slicing the image into a plurality of two-dimensional slices, each of the slices being acquired at a different scan center angle and offset by a different number of pixels from each other slice; 10. The method of any one of items 1 to 9, wherein inputting the image into the model comprises inputting the slice into the model. (Item 11) 11. The method according to any one of items 1 to 10, wherein the model is a deep learning model. (Item 12) 12. The method of any one of items 1 to 11, wherein the model is or comprises a convolutional neural network. (Item 13) 13. The method of any one of items 1 to 12, wherein the model uses a set of convolution kernels. (Item 14) 14. The method according to any one of items 1 to 13, wherein the model comprises a ResNet model. (Item 15) 15. The method according to any one of items 1 to 14, wherein the model is an Inception model. (Item 16) 16. The method of any one of items 1 to 15, wherein the predicted visual acuity metric corresponds to the predicted visual acuity of the eye of the subject. (Item 17) 17. The method of any one of items 1 to 16, wherein the predicted visual acuity metric corresponds to a predicted corrected visual acuity that predicts the visual acuity of the eye of the subject while wearing glasses or contacts. (Item 18) 18. The method of any one of items 1 to 17, wherein the predicted visual acuity metric corresponds to the predicted best corrected visual acuity of the eye of the subject. (Item 19) 18. The method of any one of items 1 to 17, wherein the subject has been previously diagnosed with age-related macular degeneration at the time the image is collected. (Item 20) 18. The method of any one of items 1 to 17, wherein the subject has been previously diagnosed with neovascular age-related macular degeneration at the time the image is collected. (Item 21) 18. The method of any one of items 1 to 17, wherein the subject has been previously diagnosed with dry age-related macular degeneration at the time the image is collected. (Item 22) 18. The method of any one of items 1 to 17, wherein each of the training subjects is a subject who was previously diagnosed with age-related macular degeneration before the training images of the set of images were collected. (Item 23) 18. The method of any one of items 1 to 17, wherein each of the training subjects is a subject who was previously diagnosed with neovascular age-related macular degeneration before the training images of the set of images were collected. (Item 24) 18. The method of any one of items 1 to 17, wherein each of the training subjects is a subject who was previously diagnosed with dry age-related macular degeneration before the training images of the set of images were collected. (Item 25) 25. The method of any one of items 1 to 24, further comprising training the model using the set of training images and the set of labels. (Item 26) 26. The method according to any one of items 1 to 25, wherein the machine learning model comprises one or more pre-processing functions and one or more neural networks. (Item 27) 26. The method of any one of items 1 to 25, wherein the image of the at least part of the eye comprises a preprocessed version of the eye generated by applying one or more preprocessing functions to a raw image of the at least part of the eye of the subject. (Item 28) 28. The method of claim 26 or 27, wherein the one or more pre-processing functions include a function that flattens the image. (Item 29) 29. The method of any one of items 26 to 28, wherein the one or more pre-processing functions include a function that generates one or more B-scan images or one or more C-scan images. (Item 30) 30. The method according to any one of items 26 to 29, wherein the one or more pre-processing functions include a trimming function. (Item 31) the predicted visual acuity metric corresponds to a predicted visual acuity of the eye of the subject at a first distance, and the method comprises: using the machine learning model to determine another predicted visual acuity metric corresponding to the predicted visual acuity of the same or a different eye of the subject; returning the other predicted visual acuity metric; and 30. The method according to any one of items 26 to 29, further comprising: (Item 32) 1. A method of stratifying a clinical trial, comprising: For each subject of the set of subjects, performing the visual acuity prediction method based on image processing described in any one of items 1 to 31; determining, for each of the set of subjects, whether the subject is eligible to participate in the clinical trial, wherein the determination is based on whether each of a set of eligibility criteria is met for the subject, and wherein evaluation of a particular eligibility criterion of the set of eligibility criteria is performed using the predicted visual acuity of the subject; conducting the clinical trial with a subset of the set of subjects, each subject in the subset being a subject determined to be eligible to participate in the clinical trial; A method comprising: (Item 33) 33. The method of claim 32, further comprising conducting the clinical trial according to the assignment of subjects to a first subject group and a second subject group. (Item 34) selecting a treatment from among a set of potential treatments for the subject based on the predicted visual acuity metric determined according to the method of any one of items 1 to 31; outputting the identification information of the treatment; A method comprising: (Item 35) 1. A method comprising: selecting a treatment from among a set of potential treatments for the subject based on the predicted visual acuity metric determined according to the method of any one of items 1 to 31; treating said subject with said treatment; A method comprising: (Item 36) Item 36. The method of item 34 or 35, wherein at least some of the training subjects are each subjects who received the treatment before participating in a vision test in which the label corresponding to the training subject is determined. (Item 37) 36. The method of claim 34 or 35, wherein each of the training subjects is a subject who received the treatment before participating in a vision test in which the label corresponding to the training subject is determined. (Item 38) Item 36. The method of item 34 or 35, wherein at least some of the training subjects each have received a treatment different from the treatment before participating in a visual acuity test in which the label corresponding to the training subject is determined, and the subjects have previously received the treatment. (Item 39) Item 36. The method of item 34 or 35, wherein each of the training subjects is a subject who has received another treatment different from the treatment before participating in the vision test in which the label corresponding to the training subject is determined, and the subject is a subject who has previously received the other treatment. (Item 40) one or more data processors; a non-transitory computer-readable storage medium containing instructions that, when executed on said one or more data processors, cause said one or more data processors to perform some or all of one or more of the methods disclosed herein; A system comprising: (Item 41) comprising instructions configured to cause one or more data processors to perform some or all of one or more of the methods disclosed herein; A computer program product tangibly embodied in a non-transitory machine-readable storage medium. [Brief explanation of the drawings]

[0008] [Figure 1] A process for predicting visual acuity of a subject's eye using a machine learning model.

[0009] [Figure 2] The process of using visual acuity prediction for clinical trials.

[0010] [Figure 3] A deep learning pipeline: prediction of best-corrected visual acuity (BCVA) by using a 3D optical coherence tomography volume represented as 30 2D images input into a ResNet-50 v2 convolutional neural network.

[0011] [Figure 4A] Actual vs. predicted best-corrected visual acuity (BCVA) at concurrent visits. Performance of a deep learning algorithm analyzing optical coherence tomography images to predict BCVA at concurrent visits. (A) Average of the test eye across all visits. (B) Average of the fellow eye across all visits. (C) Average of both the test eye and fellow eye across all visits. [Figure 4B]Actual vs. predicted best-corrected visual acuity (BCVA) at concurrent visits. Performance of a deep learning algorithm analyzing optical coherence tomography images to predict BCVA at concurrent visits. (A) Average of the test eye across all visits. (B) Average of the fellow eye across all visits. (C) Average of both the test eye and fellow eye across all visits. [Figure 4C] Actual vs. predicted best-corrected visual acuity (BCVA) at concurrent visits. Performance of a deep learning algorithm analyzing optical coherence tomography images to predict BCVA at concurrent visits. (A) Average of the test eye across all visits. (B) Average of the fellow eye across all visits. (C) Average of both the test eye and fellow eye across all visits.

[0012] [Figure 5A] Performance of a deep learning algorithm to predict best-corrected visual acuity (BCVA) <69 letters at concurrent visits from associated optical coherence tomography. (A) BCVA <69 letters. Study eye at randomized visit. (B) BCVA <69 letters. Fellow eye at randomized visit. (C) BCVA <69 letters. Both study and fellow eyes at randomized visit. [Figure 5B] Performance of a deep learning algorithm to predict best-corrected visual acuity (BCVA) <69 letters at concurrent visits from associated optical coherence tomography. (A) BCVA <69 letters. Study eye at randomized visit. (B) BCVA <69 letters. Fellow eye at randomized visit. (C) BCVA <69 letters. Both study and fellow eyes at randomized visit. [Figure 5C] Performance of a deep learning algorithm to predict best-corrected visual acuity (BCVA) <69 letters at concurrent visits from associated optical coherence tomography. (A) BCVA <69 letters. Study eye at randomized visit. (B) BCVA <69 letters. Fellow eye at randomized visit. (C) BCVA <69 letters. Both study and fellow eyes at randomized visit.

[0013] [Figure 6A]Actual vs. predicted BCVA at 12 months. Performance of a deep learning algorithm analyzing baseline optical coherence tomography images to predict BCVA at 12 months. (A) Study eye. (B) Fellow eye. (C) Both study and fellow eyes. [Figure 6B] Actual vs. predicted BCVA at 12 months. Performance of a deep learning algorithm analyzing baseline optical coherence tomography images to predict BCVA at 12 months. (A) Study eye. (B) Fellow eye. (C) Both study and fellow eyes. [Figure 6C] Actual vs. predicted BCVA at 12 months. Performance of a deep learning algorithm analyzing baseline optical coherence tomography images to predict BCVA at 12 months. (A) Study eye. (B) Fellow eye. (C) Both study and fellow eyes.

[0014] [Figure 7A] Performance of a deep learning algorithm to predict best-corrected visual acuity (BCVA) <69 letters at 12 months from baseline optical coherence tomography. (A) BCVA <69 letters at 12 months. Study eye. (B) BCVA <69 letters at 12 months. Fellow eye. (C) BCVA <69 letters at 12 months. Both study and fellow eyes. [Figure 7B] Performance of a deep learning algorithm to predict best-corrected visual acuity (BCVA) <69 letters at 12 months from baseline optical coherence tomography. (A) BCVA <69 letters at 12 months. Study eye. (B) BCVA <69 letters at 12 months. Fellow eye. (C) BCVA <69 letters at 12 months. Both study and fellow eyes. [Figure 7C] Performance of a deep learning algorithm to predict best-corrected visual acuity (BCVA) <69 letters at 12 months from baseline optical coherence tomography. (A) BCVA <69 letters at 12 months. Study eye. (B) BCVA <69 letters at 12 months. Fellow eye. (C) BCVA <69 letters at 12 months. Both study and fellow eyes.

[0015] [Figure 8] The limit of the standard deviation of BCVA in the fellow eye to the standard deviation of the test eye at baseline.

[0016] [Figure 9] R2 due to the change in BCVA for the study eye (black) and fellow eye (red) at baseline, 6 months, 12 months, 18 months, and 24 months.

[0017] [Figure 10] A deep learning pipeline for predicting best-corrected visual acuity (BCVA) using color fundus photograph images input into an Inception ResNet-v2 convolutional neural network.

[0018] [Figure 11A] Performance of a deep learning regression model analyzing color fundus photographs to predict BCVA on the ANCHOR external validation test set. (A) Actual vs. predicted BCVA at a chart distance of 2 meters. (B) Actual vs. predicted BCVA at a chart distance of 4 meters. [Figure 11B] Performance of a deep learning regression model analyzing color fundus photographs to predict BCVA on the ANCHOR external validation test set. (A) Actual vs. predicted BCVA at a chart distance of 2 meters. (B) Actual vs. predicted BCVA at a chart distance of 4 meters.

[0019] [Figure 12A] Performance of a deep learning classification model analyzing color fundus photographs to predict BCVA <69 letters (Snellen equivalent <20 / 40) on the ANCHOR external validation test set. (A) Receiver operating characteristic curves at a chart distance of 2 meters. (B) Receiver operating characteristic curves at a chart distance of 4 meters. [Figure 12B]Performance of a deep learning classification model analyzing color fundus photographs to predict BCVA <69 letters (Snellen equivalent <20 / 40) on the ANCHOR external validation test set. (A) Receiver operating characteristic curves at a chart distance of 2 meters. (B) Receiver operating characteristic curves at a chart distance of 4 meters.

[0020] [Figure 13] A network of computing systems that can be configured to perform some or all of one or more of the operations and / or one or more of the methods disclosed herein. DETAILED DESCRIPTION OF THE INVENTION

[0021] In the accompanying drawings, similar components and / or features may have the same reference label. Furthermore, various components of the same type may be distinguished by following the reference label with a dash and a second label that distinguishes between the similar components. When only a first reference label is used in this specification, the description is applicable to any of the similar components having the same first reference label, regardless of the second reference label.

[0022] I. Overview This description relates to predicting a metric characterizing a subject's visual function (e.g., visual acuity) based on an analysis of an image of the subject's eye. The prediction can facilitate the definition of objectives during treatment development and the selection of an appropriate treatment for a given individual. The image of the eye can include (for example) an optical coherence tomography (OCT) image, a color fundus image, or an infrared fundus image. The subject can be, but need not be, diagnosed with an ocular disease (e.g., macular degeneration, neovascular age-related macular degeneration, glaucoma, cataracts, or diabetic retinopathy).

[0023] The predicted visual acuity can include a visual acuity metric generated based on how accurately the subject can identify visual objects (e.g., of one or more sizes). The visual objects can include alphanumeric characters, and / or the visual acuity can correspond to the ability to read one or more alphanumeric characters on an eye chart. The predicted visual acuity can correspond to a metric corresponding to current visual acuity (e.g., characterizing visual acuity at a time or time period related to the capture of the eye image) or future visual acuity (e.g., characterizing visual acuity at a future time or time period compared to the visual acuity at or during the time the eye image was captured).

[0024] Visual acuity can be predicted (for healthy, disease-free subjects, for subjects with eye-related diseases, or for subjects with eye-related pathologies) using a trained machine learning model. In some cases, the machine learning model may be configured to output the results of a logic-based condition evaluation. For example, the machine learning model can be configured and trained to predict whether an observer's current or future visual acuity (e.g., at a particular time) will be below a predetermined threshold. The machine learning model can include (for example) a deep machine learning model, a convolutional machine learning model, a regression model, and / or a classifier model. Exemplary deep convolutional neural network architectures that can be used include AlexNet, ResNet, VGGNet, Inception, DenseNet, EfficientNet, GoogLeNet, or variations of any of these models. The machine learning model can include (for example) one or more convolutional layers, one or more pooling layers (e.g., max pooling or average pooling), one or more fully connected layers, one or more dropout layers, and / or one or more activation functions (e.g., a sigmoid function, a tanh function, or a rectified linear unit). The machine learning model can use filters of one or more kernel sizes in the convolutional layers of the model. For example, the kernel sizes can vary between layers. The machine learning model can be configured and trained to output (for example) a predicted visual acuity metric, a visual acuity metric range (e.g., open or closed range), and / or a result based on an evaluation of one or more logical conditions (e.g., to indicate whether the subject's visual acuity is at least as good as a particular threshold).

[0025] A trained machine learning model can be defined based on a set of fixed hyperparameters (e.g., defined by a programmer or user) and a set of parameters (e.g., having values ​​learned during training). The set of parameters can include (for example) a set of weights that a node (or neuron) applies to transform received values. The values ​​received by a given node can include a set of values ​​from a previous layer (or from an input), where the set of values ​​corresponds to the node's received field (e.g., where the received field has a size, padding, etc. defined by one or more hyperparameters).

[0026] In some cases, the machine learning model is trained to receive (e.g., partially or entirely) inputs including data corresponding to a particular time point or time period and output results corresponding to the particular time point or time period. For example, the machine learning model can be trained to predict a subject's current visual acuity or the visual acuity at the date one or more input images were originally acquired. In some cases, the machine learning model is trained to receive (e.g., partially or entirely) inputs including data corresponding to a particular time point or time period and output results corresponding to a later time point or time period (e.g., at least one week from the date the input data was collected, at least two weeks from the date the input data was collected, at least one month from the date the input data was collected, at least three months from the date the input data was collected, at least six months from the date the input data was collected, at least one year from the date the input data was collected, at least two years from the date the input data was collected, or at least six years from the date the input data was collected).

[0027] In some cases, the machine learning model is initialized with parameters learned during training to predict a current metric (e.g., corresponding to a time or time period associated with the input data) and is then further trained (e.g., using transfer learning) to predict a metric for a later point in time or time period (e.g., at least one week from the date the input data was collected, at least two weeks from the date the input data was collected, at least one month from the date the input data was collected, at least three months from the date the input data was collected, at least six months from the date the input data was collected, at least one year from the date the input data was collected, at least two years from the date the input data was collected, or at least six years from the date the input data was collected). In some cases, the machine learning model is configured to receive input values ​​indicating a subsequent time point at which a prediction will be made relative to a time associated with the set of images included in the input (e.g., three months, six months, one year, etc.). In some cases, one or more machine learning models are separately trained to predict outputs associated with different points in time or time periods.

[0028] In some cases, the processing pipeline includes a machine learning model for processing one or more input images to generate an intermediate result (e.g., a prediction of visual acuity or best-corrected visual acuity associated with a current time point or a time point associated with the set of input images), and further includes one or more other machine learning models and / or one or more other post-processing functions configured to receive the intermediate result (e.g., representative of predicted visual acuity at one or more first time points) and output another result (e.g., representative of predicted visual acuity at a second time point subsequent to the first time point). For example, the another machine learning model and / or post-processing function may include a regression model configured to receive the intermediate result and potentially one or more other intermediate results related to one or more other variables (e.g., one or more other time points, one or more other subject characteristics, one or more other subject-specific ocular anatomical structure values, one or more variables indicative of a disease or disease state, etc.) and output a final (e.g., subject- and / or eye-specific visual acuity metric).

[0029] The one or more predicted visual acuities can be used to inform (for example) diagnosis, prognosis, treatment selection, treatment recommendation, dosage selection, and / or dosage recommendation. The one or more predicted visual acuities may be associated with the same subject, may be associated with the time or period when images of the subject's eye were collected, and / or may be associated with a time or period after the time or period when images of the subject's eye were collected. In some cases, a proportional or absolute difference between the predicted visual acuity for a future time point and the predicted visual acuity for a current or previous time point is determined and used to inform diagnosis, prognosis, clinical trial criteria, treatment selection, treatment recommendation, dosage selection, dosage recommendation, recommended frequency of follow-up by a medical professional, etc. Diagnosis can include (for example) identifying a disease (e.g., age-related macular degeneration), a type of disease (e.g., neovascular age-related macular degeneration vs. atrophic age-related macular degeneration), or disease severity (e.g., severity of age-related macular degeneration). Prognosis can include identifying (for example) the predicted time for a given decline in visual function, the predicted time for a given disease transition, the probability of a given disease transition occurring (e.g., over a period of time or at any time), etc. For example, approximately 10-20% of subjects with dry AMD transition to neovascular AMD. Predicted visual acuity relative to a future time can be used to predict the probability or predicted time for a subject to transition to neovascular AMD.

[0030] As an example, if a subject has been diagnosed with age-related macular degeneration, the disease can be "neovascular age-related macular degeneration (AMD)" (also known as wet AMD) or "dry AMD" (also known as dry AMD). Approximately 80-90% of individuals with AMD have dry AMD. Neovascular AMD is more aggressive and accounts for approximately 90% of severe, vision-loss AMD cases. In some cases (e.g., if the data indicates that the subject has not been specifically diagnosed with neovascular AMD or dry AMD), the output can predict that the subject is likely to have neovascular AMD (instead of dry AMD) if the relative or absolute difference between predicted future visual acuity and current or recent visual acuity exceeds a predetermined change threshold, or if the absolute predicted future visual acuity exceeds a predetermined present-time threshold. In some cases (e.g., when an early diagnosis of dry AMD is indicated or there are no reports of vascular fluid leakage), if the relative or absolute difference between predicted future visual acuity and current or recent visual acuity exceeds a predetermined change threshold, or if the absolute predicted future visual acuity exceeds a predetermined current time threshold, the output may predict that the subject is likely to have late dry AMD (instead of early dry AMD). In some cases (e.g., when an early diagnosis of dry AMD is indicated or there are no reports of vascular fluid leakage), if the relative or absolute difference between predicted future visual acuity and current or recent visual acuity exceeds a predetermined change threshold, or if the absolute predicted future visual acuity exceeds a predetermined current time threshold, the output may predict that the subject is at relative risk of transitioning to neovascular AMD (e.g., within a given time period).

[0031] Most subjects with dry AMD are not treated with medication, but instead are monitored to determine whether and / or if their condition will progress to neovascular AMD. If the output predicts that the subject is likely to have neovascular AMD, likely to have late dry AMD, and / or at high risk of progressing to neovascular AMD, the computing system can recommend and / or a caregiver can decide to reschedule the subject's image-based monitoring (e.g., so that eye images are collected earlier than previously planned) and / or initiate AMD treatment. AMD treatments can include (for example) anti-vascular endothelial growth factor (anti-VEGF) agents (e.g., ranibizumab, bevacizumab, afibercept, and conbercept) or faricimab (a bispecific antibody currently in clinical trials that neutralizes angiopoietin-2 and VEGF-A).

[0032] Glaucoma is another eye disease that causes vision impairment (e.g., vision loss). Glaucoma is caused by abnormally high pressure in the eye, which damages the optic nerve. Risk factors for glaucoma include age, race, family history, and eye injury. People over the age of 60 are at higher risk of developing glaucoma. African Americans are more likely to develop glaucoma, and those over the age of 40 are at higher risk. Additionally, medical conditions such as diabetes and high blood pressure can increase the risk of developing glaucoma.

[0033] Glaucoma can be diagnosed through a comprehensive eye examination. Visual acuity testing and intraocular pressure measurement can be used to facilitate the diagnosis of glaucoma. Additionally, OCT and color fundus photography can identify changes or abnormalities in the optic nerve that may indicate glaucoma.

[0034] Glaucoma develops slowly but can cause blindness over 20 years if left untreated. Treatment can prevent blindness. In less severe cases, prescription eye drops can be used to make the eye produce less fluid. In more severe cases, laser surgery can be used to open up the eye's drainage network. Imaging techniques such as OCT and color fundus photography can be used to monitor the development of glaucoma over time. The rate of progression can be useful in selecting treatment options.

[0035] Other common age-related eye conditions that affect vision are cataracts and dry eye. Cataracts are clouding of the eye's lens, and dry eye is a lack of proper lubrication within the eye. Cataracts can be surgically removed, and dry eye can be treated with eye drops or other medications. In extreme cases, surgical procedures can be performed to treat dry eye. Normal vision or BCVA can be restored by treating both cataracts and dry eye.

[0036] Another ocular condition that can cause vision loss is diabetic retinopathy. In the early stages of the disease, mild, non-proliferative abnormalities occur, such as increased vascular permeability through the blood vessel walls (increasing the flow of small molecules or whole cells). In later stages, vascular occlusion and / or new blood vessel growth on the retina or posterior vitreous are frequently observed. Macular edema can occur at any stage of the disease and involves the accumulation of fluid in the layers of the macula due to ruptured blood vessels. Macular edema can result in blurred vision and decreased vision.

[0037] All diabetics are at risk for developing diabetic retinopathy. Diagnosis can be made by detecting (for example) decreased vision, microaneurysms (e.g., as shown in fundus photographs or optical coherence tomography (OCT) images), dot and blot hemorrhages (e.g., as shown in fundus photographs or OCT images), hard retinal exudates (e.g., as shown in fundus photographs), soft exudates (e.g., as shown in fundus photographs), venous dilation (e.g., as shown in fundus photographs or OCT images), retinal thickening (e.g., as shown in fundus photographs or OCT images), leakage or nonperfusion of the retinal and / or choroidal vasculature (e.g., as shown by the use of fundus fluorescein angiography), etc. Diabetic retinopathy is further consistent with multiple pupil size measurements (e.g., observed after pupil dilation), such as a decreased baseline pupil diameter, decreased pupil constriction amplitude, decreased pupil constriction rate, and / or decreased pupil dilation rate.

[0038] Early diabetic retinopathy is often left untreated. Advanced diabetic retinopathy can be treated with anti-VEGF drugs, vitrectomy (surgery to remove blood and scar tissue from the eye), panretinal photocoagulation (laser treatment to shrink blood vessels) and / or photocoagulation (laser treatment to prevent leakage of blood and other fluids into the eye).

[0039] The techniques disclosed herein can be used to process images of the eyes of subjects with an ocular disease or condition (e.g., age-related macular degeneration, diabetic retinopathy, macular edema, glaucoma, cataract, or dry eye) to predict (for example) the effectiveness of a particular treatment, identify a particular treatment to use or recommend and / or facilitate the design or conduct of a clinical trial.

[0040] II. Input data, preprocessing, and training data II.A. Input Data Data received by a machine learning model (e.g., to be processed to generate new outputs or to be processed during training) can include one or more images of the subject's eye, one or more processed versions of which can be captured using one or more imaging techniques, such as optical coherence tomography (OCT), color fundus photography, fundus autofluorescence, or infrared fundus photography. The one or more imaging techniques can be non-invasive and / or can be non-administered with a dye (e.g., intravenously or orally). In other examples, the one or more imaging techniques can include techniques that involve administering a dye (e.g., as done in fundus fluorescein angiography).

[0041] In some cases, the input dataset includes a single image. In some cases, the input dataset includes a set of images, where the images in the set correspond to different depths.

[0042] II.A.1. Optical Coherence Tomography OCT is a noninvasive imaging technique that uses light waves to construct cross-sectional images of the eye. To generate an image of the eye, an interferometer can split a low-coherence near-infrared light beam toward the target tissue and a reference mirror. The backscattered light received from the target tissue and the reference mirror is combined, and the signal interference is determined. Areas of the target tissue that reflect more light have higher interference. The reflectance information, often referred to as an A-scan, contains information about the longitudinal axis (e.g., depth) of the target tissue. The light beam can be guided in a linear direction to generate information at many lateral positions. Cross-sectional "B-scan" images can be generated by combining depth scans at lateral positions (e.g., using 128,256, or 512 A-scans). Cross-sectional images can be displayed in real time, speeding up the analysis and diagnosis process. Three-dimensional images can be constructed to include multiple B-scans.

[0043] OCT can provide high-resolution images of the eye without contact with the eye. OCT allows distinct layers of the retina to be examined, which can be useful in identifying diseases originating in specific layers. Two potentially important areas of the retina include the optic nerve and the macula. The optic nerve carries information from the eye to the brain, and the macula is an area with a high density of photoreceptor cells.

[0044] OCT can provide information about the thickness and size of various layers of the eye, their cellular organization, and even the thickness of axons. As a result, OCT can aid in the diagnosis and treatment of conditions such as neovascular AMD. OCT images depicting or suggesting fluid in the retinal, subretinal, or subpigmentary epithelial space (e.g., dome-shaped reflective areas in the subretinal space, elevated non-drusenoid retinal pigment epithelium or drusenoid pigment epithelial detachment, and retinal thickening) can be consistent with the presence of neovascular AMD.

[0045] II.A.2. Color fundus photography Color fundus photography involves recording color images of the inner surface of the eye using a fundus camera or retinal camera. A fundus camera is a low-power microscope equipped with a camera that can capture images of the retina, retinal vasculature, optic disc, macula, and fundus of the eye. The light beam (e.g., white light) sent from the fundus camera to capture the image passes through the pupil of the eye. In some cases, the pupil of the eye can be dilated to provide a larger area that can be captured. Color fundus images show the retina, macula, blood vessels, optic disc, and fundus.

[0046] Drusen (mainly subretinal lipid deposits) fluoresce and appear white or yellow in color fundus images. Drusen are commonly detected in older adults, but the presence of large and / or massive drusen in the macula is frequently observed in subjects with AMD.

[0047] II.B. Pretreatment As described above, raw images can be processed to generate other images having different perspectives and / or dimensions relative to the raw image. For example, multiple A-scans can be processed to generate one or more B-scans, and multiple B-scans can be used to generate a C-scan (which can capture some three-dimensional information, such as depth). In some cases, the image can include a three-dimensional image (e.g., generated based on multiple two-dimensional images). In some cases, color fundus photographs can be collected at different imaging angles, facilitating the generation of images that can convey depth information.

[0048] Due to the curvature of the eye, the raw image may depict curved structures (e.g., a curved retinal layer). A flattening technique may then be used to flatten an image (e.g., a two-dimensional or three-dimensional image) based on (for example) a pilot estimate of a given structure (e.g., the retinal pigment epithelium). For example, the image may be filtered (e.g., using a Gaussian filter) to remove noise from the image, and the most intense pixel in each column of the denoised image may be identified (e.g., which may represent the retinal pigment epithelium surface). The columns may then be adjusted up and down to align the most intense pixels across the columns. As another example, the retinal pigment epithelium surface may be segmented (e.g., via intensity thresholding), and a function (e.g., a spline function) may be fitted to the segmented pixels. The columns of the image may then be realigned to flatten the spline function. In one example, the raw image may be flattened to a segmentation of the retinal pigment epithelium layer, and the image volume may be cropped to pixels above and below the flattened retinal pigment epithelium.

[0049] Pre-processing can include normalizing and / or standardizing the intensity values, which can be performed before or after other types of pre-processing (e.g., before or after generating a B-scan, generating a C-scan, or applying a flattening technique).

[0050] Flattening and / or other preprocessing techniques can be performed to align depictions of particular structures to target locations. For example, pixels suspected to correspond to the retinal pigment epithelium surface can be shifted (e.g., during the flattening process) to a specified row or plane so that the same row or plane corresponds to the same structure across the image.

[0051] Some machine learning models can be configured to receive images of a particular size. Thus, preprocessing can be performed, for example, to crop and / or pad the image so that its dimensions meet a particular size. Some machine learning models can be configured to receive images with a particular resolution. Thus, preprocessing can be performed, for example, to downsample or upsample the image.

[0052] II.C. Training Data A training data set includes a plurality of training data elements, each associated with a particular eye of a particular subject. Each of the plurality of training data elements can be further associated with a particular time point and / or medical visit. In some cases, a given training data set corresponds to a particular type of disease. For example, a training data set can be defined to correspond to a set of subjects each having AMD or each having neovascular AMD. In some cases, one or more other constraints are imposed on the subject set (e.g., all subjects in the subject set are within a particular age range, are not taking medication, etc.).

[0053] Each of the plurality of training data elements can include input data of one or more images of at least a portion of an eye. In some cases, the images processed by the machine learning model are preprocessed versions of the raw images (e.g., preprocessed using preprocessing techniques such as those disclosed in Section II.B.). In some cases, the machine learning model includes one or more preprocessing functions for preprocessing the received images (e.g., including one or more of those disclosed in Section II.B.).

[0054] Each of the plurality of training data elements can further include a label including or otherwise indicating the visual acuity associated with a particular subject and associated with a particular time point, where the label can indicate the degree to which the subject discriminates between visual stimuli. The visual acuity metric can be determined by determining whether and / or the degree to which a particular subject can accurately identify and / or characterize one or more visual stimuli presented to the subject (e.g., at a particular distance from the subject and / or having one or more particular sizes). The visual acuity metric can be specific to a particular eye of a subject by evaluating the response provided when the particular subject views one or more visual stimuli with only that particular eye (e.g., with the other eye blocked or closed).

[0055] II.C.1. Types of visual acuity metrics The predicted visual acuity metric can include a numerical metric such as (for example) a ratio, a numerator (e.g., relative to a fixed denominator), a real number, an integer, etc. The predicted visual acuity metric can include a visual acuity category and / or visual acuity boundaries. For example, a visual acuity scale can include a set of threshold visual acuities (e.g., 20 / 10, 20 / 20, 20 / 25, 20 / 30, 20 / 40, 20 / 50, 20 / 70, 20 / 100, and 20 / 200 or −0.3, −0.2, −0.1, 0, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, and 0.7), each associated with one or more visual stimuli such that if the observer can correctly identify the visual stimulus (or a characteristic thereof), the observer has at least the associated threshold visual acuity. The predicted visual acuity metric can be defined as the "best" across the set corresponding to results from a particular user (e.g., indicating that the observer can accurately identify or characterize the stimuli associated with the visual acuity metric, but cannot identify or characterize other stimuli associated with higher visual acuities in the set). In some cases (e.g., when using the Snellen scale), a higher visual acuity metric represents better visual acuity compared to other metrics associated with lower visual acuity metrics. In some cases (e.g., when using the LogMAR scale), a lower visual acuity metric represents better visual acuity compared to other metrics associated with higher visual acuity metrics. In some cases (e.g., when using the Jaeger scale), values ​​along the scale are not monotonically dependent on visual acuity (e.g., because J1+ represents better visual acuity than J1, and J1 represents better visual acuity than J2).

[0056] The visual acuity metric can include a value selected from among values ​​expressed along a scale such as the Snellen scale, the LogMAR scale, or the Jaeger scale. The visual acuity metric can include a Snellen fraction, a LogMAR value, or a Jaeger score. The visual acuity metric can include a value determined based on the subject's ability to correctly identify and / or characterize one or more characters (e.g., when viewing the one or more characters with a particular eye). The one or more characters can include one or more letters, one or more numbers, one or more tumbling Es, one or more Landolt Cs, and / or one or more Lear symbols presented at a particular size and a particular distance from the subject.

[0057] A visual acuity metric can include a value determined based on an observer's ability to identify and / or characterize one or more visual stimuli presented on a chart or card placed at a specific distance (e.g., 2 meters or 4 meters) from the observer. The chart or card can be a physical chart or card (e.g., comprising paper, plastic, laminate, cardboard, etc.) or a virtual chart or card (e.g., presented on the display of an electronic device). The chart or card can include multiple lines, each containing one or more characters (e.g., letters, numbers, Tumbling E, Landolt C, and / or Lear symbols). Each line can be associated with a character size (e.g., height, width, and / or aspect ratio) indicating the size of one or more characters within the line. The size can vary (e.g., monotonically) across the line. The number of lines on the chart or card can be (for example) at least 3, at least 5, at least 7, at least 9, or at least 11 lines, fewer than 20, fewer than 15, fewer than 13, fewer than 10, or fewer than 7 lines, and / or about 5, about 7, about 9, and / or about 11 lines.

[0058] Charts or cards can include (for example) Snellen charts (e.g., taller than wide, containing multiple lines of letters, each line having more letters and smaller letters than the previous line, and the letter size difference between lines varying across the chart, considering both absolute and relative size differences), modified Snellen charts, logarithm of minimum angle of resolution (LogMAR) charts (same number of letters per line, letter size, and line spacing varying logarithmically across lines and letters sized in a square shape), Bailey-Loewy charts, Early Treatment Diabetic Retinopathy Study (ETDRS) charts (same number of letters per line, letter size, and line spacing varying logarithmically across lines and letters sized in a rectangular shape), Rosenbaum cards.

[0059] The visual acuity metric can include visual acuity at a particular distance (e.g., 2 meters or 4 meters) with or without vision correction (e.g., glasses or contacts). If the visual acuity metric is determined when the viewer is using vision correction, the metric can include corrected visual acuity or best-corrected visual acuity.

[0060] Visual acuity metrics can include ratios. A ratio can include a numerator representing an estimated low-threshold distance (in a particular unit) that an observer must be from a visual stimulus to decode it as easily as an unimpaired observer can decode it at a distance (in a particular unit), as indicated in the ratio's denominator. A ratio can include a numerator identifying the maximum distance a subject can be from a visual stimulus to accurately identify or characterize the visual stimulus (e.g., using one eye), and a denominator identifying the maximum distance a representative subject (without visual impairment) can be from the visual stimulus (or a similar visual stimulus having the same or substantially similar height and / or width as the visual stimulus) to accurately identify or characterize the visual stimulus. A ratio can include a Snellen fraction. Ratios and / or tests can be defined to produce ratios with the same value for the numerator or denominator. As an example, a 20 / 40 metric can indicate that an observer must be within 20 measurements (e.g., feet, yards, meters, etc.) of a visual stimulus to detect and / or decode stimulus features that an average unimpaired observer can detect and / or decode in 40 measurements. Ratios can have a specified numerator. For example, a model can be configured to output a ratio with the numerator set to 20.

[0061] In some cases, all of the training data elements include the same type of visual acuity metric. For example, all of the training data elements may include Snellen fractions. In some cases, at least two of the training data elements include different types of visual acuity metrics. For example, a first training data element may include a Snellen fraction and a second training data element may include a LogMAR score. In these examples, each of at least some of the training data elements may be processed to convert the visual acuity metric. For example, a lookup table or algorithm may be used to convert a LogMAR score to a Snellen fraction, or a lookup table or algorithm may be used to convert each LogMAR score and each Snellen fraction to a score along yet another scale.

[0062] II.C.2. Timing Associations for Visual Acuity Metrics In some cases, for each of one or more of the training data elements, the visual acuity metric includes one determined based on a functional assessment of the subject's eye (e.g., by answers provided in response to presentation of a chart or card) performed at a time corresponding to the date the image in the training data element was collected. For example, the subject may complete a visual acuity assessment (e.g., by attempting to correctly identify or characterize letters or other visual stimuli on an eye chart) on the same day, within the same week, or within the same two weeks as the image of the subject's eye was collected.

[0063] In some cases, for each of the one or more training data elements, the visual acuity metric includes one determined based on a functional assessment of the subject's eye performed at a time after the date the images in the training data element were collected. For example, the subject may complete a visual acuity assessment at least or about one month, at least or about two months, at least or about six months, at least or about one year, at least or about two years, or at least or about five years after the date or period the images of the subject's eye were collected.

[0064] In some cases, for each of one or more of the training data elements, the training data element includes a visual acuity metric associated with a time corresponding to the date the images in the training data element were collected, and further includes another visual acuity metric associated with a time after the date the images in the training data element were collected.

[0065] III. Machine Learning Models The training data can be used to train a machine learning model so that a set of parameters is learned. The trained machine learning model can then be used to process other inputs (e.g., of the types described in Section II.A, which can be preprocessed using types such as those disclosed in Section II.B) and generate an output that predicts visual acuity.

[0066] III.A. Model Architecture The machine learning model may have (for example) a deep learning architecture, a residual network architecture, and / or a convolutional neural network architecture. The deep learning architecture may be configured such that the model performs hierarchical learning, first identifying low-level features at one or more lower levels and then identifying high-level features at a higher level. The machine learning model may have (for example) a ResNet architecture, an AlexNet architecture, a DenseNet architecture, an EfficientNet architecture, a GoogLeNet architecture, or a VGGNet architecture.

[0067] A machine learning model can include one or more convolutional layers (e.g., at least three convolutional layers, at least four convolutional layers, at least five convolutional layers, or at least seven convolutional layers), one or more sparse convolutional layers, one or more pooling layers (e.g., max pooling or average pooling layers), one or more initiation modules, the use of dropout, the use of batch normalization, one or more dense layers, and / or one or more activation functions. In some cases, one or more or all of the convolutional layers are each followed by a pooling layer. For example, a model can include five convolutional layers, and three convolutional layers can each be followed by a pooling layer. The pooling layers can reduce the number of parameters learned by the model.

[0068] A machine learning model may include a residual layer or skip connection that feeds the output from one layer (layer l) to another layer that is not directly adjacent to the layer (layer l+2, layer l+3, layer l+4, etc.). Skipping can reduce or avoid vanishing gradient situations. Thus, skipping can further facilitate the use of networks with a large number of layers. A machine learning model may include (for example) at least 10, at least 20, at least 30, at least 40, or at least 45 convolutional layers (e.g., potentially one or more max-pooling layers, as well as one or more average pooling layers and / or one or more fully connected layers). A machine learning model may include residual blocks, each including three layers. The size of the kernel used in the residual block in an earlier stage of the model may be smaller than the size of the kernel used in the residual block in a later stage of the model.

[0069] The machine learning model can be configured to output a numeric output or a categorical output (e.g., indicating a level of visual acuity and / or including a binary prediction as to whether visual acuity is above a threshold). In some cases, a machine learning model that outputs a numerical output includes a linear activation function and / or uses an error-based loss function (e.g., mean squared error). A machine learning model that outputs a categorical output can include (for example) a softmax activation function and / or use a categorical loss function (e.g., a cross-entropy loss function for sparse categories).

[0070] III.B. Model Training In some cases, the first training data set includes a first set of training data elements, each of which includes one or more images (of a subject's eye) and a label corresponding to the subject's visual acuity at a time corresponding to the time the image was captured. A machine learning model having a particular architecture can be trained using the first training data set such that a first set of parameters is learned.

[0071] The second training dataset can include a second set of training data elements, each of which includes one or more images (of a subject's eye) and a label corresponding to the subject's visual acuity at a time later than the time the image was captured (e.g., 12 months later). In some cases, each subject represented in the first training dataset is also represented in the second training dataset. In some cases, at least some of the subjects represented in the first training dataset are also represented in the second training dataset. In some cases, the subjects represented in the first training dataset are at least partially or completely different from the subjects represented in the second training dataset.

[0072] A machine learning model having an architecture at least partially or entirely the same as the particular architecture can be trained using a second training dataset such that a second set of parameters is learned. For example, each of the models trained using the first and second training datasets can include a ResNet-50 CNN architecture.

[0073] In some cases, a model to be trained to predict subsequent visual acuity metrics is initialized using parameters learned by a model trained to predict current visual acuity (e.g., corresponding to the same time period during which the input images were collected), after which subsequent training can be performed.

[0074] In some cases, a single machine learning model may be trained using both the first and second training datasets, but the machine learning model may be configured to receive instructions (in addition to the image) regarding to which point in time the prediction corresponds.

[0075] In some cases, the machine learning model can be configured to predict the rate of visual acuity decline, one or more rate constants, etc. to predict how the visual acuity of the subject's eye will change over time (e.g., how quickly the visual acuity will decline). The predictions can include constants used in (e.g.,) linear, logarithmic, polynomial, and / or exponential functions to characterize the visual acuity over time.

[0076] IV. Use of Model Output for Clinical Trial Design, Prognosis, and Treatment Selection Model outputs that predict current visual acuity metrics, predict visual acuity metrics at specific future time points, and / or predict how visual acuity will change can be used (for example) to inform clinical trial design, subject prognosis, and treatment selection.

[0077] IV.A. Clinical Trial Design In some cases, clinical trials can be defined such that eligibility criteria relate to observed or predicted visual acuity. For example, a clinical trial can be defined to enroll subjects with visual acuity of 20 / 80 or better (and meeting other eligibility criteria). Visual acuity-based eligibility criteria can be defined as relating to a visual acuity metric generated based on images of at least a portion of the subject's eye, characterizing current or past visual acuity. The criteria may require a predicted metric or require acceptance of either a predicted or observed metric. Alternatively, or in addition, visual acuity-based eligibility criteria may be defined as relating to a predicted visual acuity metric characterizing predicted visual acuity at a later time point (e.g., requiring predicted visual acuity one year from the current time to be worse than 20 / 40).

[0078] Many clinical trials (e.g., clinical trials) are controlled so that one group of subjects receives an investigational treatment, while a second group of subjects receives a different treatment, a control treatment, no treatment, and / or standard care. Model results corresponding to predicted visual acuity can be used to facilitate the selection of a group of subjects such that the first group of subjects resembles the second group of subjects. Currently, there is considerable variability across AMD subjects in terms of how rapidly visual function declines. Thus, two AMD subjects with identical visual acuity at one time may have significantly different visual function the second time. This can lead to unintentionally defining groups of subjects so that one group progresses much faster than the other, potentially confounding results.

[0079] Therefore, clinical trial stratification can be designed to include eligibility criteria that specify a range of predicted visual acuity at a future time point. For example, a clinical trial may require that OCT or color fundus images be collected within a first time period (before the trial begins) and that the predicted visual acuity at a second time point (e.g., 12 or 18 months after the first time period) be within a predetermined range. The criteria can further identify constraints (e.g., ranges) of predicted visual acuity that correspond to the time the images were collected.

[0080] Additionally or alternatively, predicted visual acuity can be used to process clinical trial data and develop indications and / or hypotheses regarding the types of subjects who are particularly likely (or particularly unlikely) to benefit from a given treatment. For example, during or after the trial, for each subject in the treatment group, current metrics such as current visual acuity (e.g., as determined using an eye chart or card), predicted current visual acuity (e.g., as determined using a current image of the eye), and current diagnosis of AMD (e.g., whether the AMD is wet or dry AMD) can be determined. Then, it can be determined whether a predicted visual acuity metric corresponding to a time before the trial (e.g., related to the period when the image was taken) predicts the current metric, or whether another predicted visual acuity metric corresponding to a subsequent time predicts the current metric.

[0081] IV.B. Prognosis The predicted visual acuity metric can be used as part of or to inform the prognosis identified for the subject. A subject can have a worse prognosis if the subject's visual acuity at a later time point (e.g., 12 months) is predicted to be worse than a threshold or correspond to a substantial decline, compared to a subject with a predicted subsequent visual acuity corresponding to a diagnosis that is better than the threshold or less substantial. The prognosis can affect the subject's appearance, treatment selection, etc.

[0082] IV.C. Treatment Selection There is substantial variability in the degree to which different subjects respond to a given treatment.For example, anti-vascular endothelial growth factor (anti-VEGF) is the standard treatment for neovascular AMD.However, there is high variability in the degree to which subjects respond.

[0083] Disease activity can be present, for example, when visual acuity declines or does not improve, when central retinal thickness observed by OCT does not decrease, when new intraretinal fluid is detected, when new subretinal fluid is detected, and / or when new retinal thickening is detected. Thus, predicted future visual acuity can indicate whether a given subject will effectively respond to a particular treatment. For example, a model can be trained using a training dataset that includes images of the subject's eye before the treatment is administered and includes (as labels) predicted visual acuity (or predicted visual acuity change) at a subsequent time point, and the subject receives a particular treatment between a first time period (related to image collection) and a second time point (related to subsequent visual acuity).

[0084] Thus, a caregiver can predict how much vision will change while using a particular treatment. Multiple models can even be used to evaluate different possible treatments. A caregiver can then predict whether a particular treatment will be effective for a particular subject and / or which of multiple treatments will be most effective for a particular subject.

[0085] For example, for neovascular AMD, one treatment option is anti-VEGF injections (here, VEGF stands for a protein called vascular endothelial growth factor, which can promote the formation of new blood vessels at the back of the eye and can cause macular degeneration due to leakage of blood and other body fluids). Anti-VEGF injections are injections into the vitreous of the eye to stop the abnormal growth of blood vessels. The most common anti-VEGF injections are ranibizumab, bevacizumab, conbercept, and aflibercept. A person's response to anti-VEGF injections is often monitored by OCT or manual evaluation of color fundus photographs to determine how often they should receive injections.

[0086] Patients may receive monthly anti-VEGF injections until it is determined that the interval between injections can be extended. Vision loss due to neovascular AMD is frequently, but not always, halted by the use of anti-VEGF injections. Occasionally, subjects recover some vision loss as a result of anti-VEGF injections.

[0087] Another treatment option (currently in Phase III clinical trials) is a port delivery system (PDS) used in combination with anti-VEGF medications. A PDS is a permanently refillable ocular implant that can continuously deliver anti-VEGF medications to the eye for several months. A PDS can allow subjects to visit an ophthalmologist only twice a year for medication refills. Reducing the number of required visits to an ophthalmologist can reduce treatment burden, which often leads to undertreatment and poor vision outcomes.

[0088] Yet another treatment option is faricimab, an antibody currently in Phase III clinical trials. Faricimab is administered intraocularly. Some vision loss can be reversed in response to treatment.

[0089] Furthermore, another treatment option for neovascular AMD is photodynamic therapy (PDT). PDT is a laser treatment that destroys excess blood vessels. If the subject experiences gradual rather than sudden vision loss, doctors can choose PDT instead of anti-VEGF injections. In some cases, PDT can be combined with anti-VEGF treatment to delay central vision loss caused by neovascular AMD.

[0090] Thus, model predictions can be used to inform decisions about whether to recommend or use an anti-VEGF agent, an anti-VEGF agent delivered via intraocular injection, a specific VEGF agent (e.g., ranibizumab or bevacizumab), PDS or PDT for a particular subject and / or subject's eye.

[0091] IV.D. Initiating Treatment and / or Monitoring Changes In some cases, the machine learning model output can be used to recommend, prescribe or administer AMD treatment to subjects who are not currently receiving AMD treatment, or in some cases, are not receiving AMD treatment.For example, a subject may have a diagnosis of dry AMD, which is often not treated with medication.However, the output from the machine learning model (for example, trained using data corresponding to subjects diagnosed with dry AMD, subjects diagnosed with neovascular AMD, or both) can correspond to a prediction that the visual acuity (for example, best-corrected visual acuity) corresponding to the baseline date (when the input image is collected), the same visit, or subsequent time (for example, 3 days, 5 days, 1 week, 2 weeks, 1 month, 6 months, 1 year, 2 years, 5 years from the baseline date) is worse than a predetermined threshold (for example, corresponding to the Snellen fraction of 20 / 30, 20 / 40, 20 / 60 or 20 / 80). Alternatively or additionally, output from the machine learning model (e.g., trained using data corresponding to subjects diagnosed with dry AMD, subjects diagnosed with neovascular AMD, or both) can correspond to a prediction that visual acuity will decrease by at least a threshold fraction or absolute predetermined amount (e.g., corresponding to a Snellen fraction at the subsequent time that is less than 90%, less than 75%, less than 66%, or less than 50% of the Snellen fraction at the baseline time).

[0092] In response to one or more of these types of predictions, a computing system (e.g., one that generated the prediction or one that received the prediction from another computing system) can output a recommendation for a specific action, a caregiver (e.g., a doctor, nurse, clinic, hospital, etc.) can provide a recommendation for a specific action, and / or a specific action can be performed. A specific action can include (for example) changing the subject's diagnosis (e.g., from dry AMD to neovascular AMD or from early dry AMD to late dry AMD), initiating the subject's AMD treatment (e.g., including one or more treatments identified herein), modifying the subject's treatment schedule (e.g., increasing the frequency of anti-VEGF administration), and / or modifying the subject's AMD treatment (e.g., to one identified herein). For example, a specific action can include initiating the subject's anti-VEGF treatment. In some cases, the particular action includes (for example) changing the monitoring schedule so that the next date for imaging the subject's eyes or vision testing is scheduled or changed (e.g., to an earlier date), increasing the frequency of imaging the subject's eyes, and / or increasing the frequency of testing the subject's vision.

[0093] V. Exemplary Processes FIG. 1 illustrates a process 100 for predicting the visual acuity of a subject's eye using a machine learning model. Process 100 begins at block 105, where the machine learning model is trained to learn a set of parameters using a set of training images and a set of labels. Each of the set of training images can depict at least a portion of a training subject's eye, such that the set of training images depicts at least a portion of each of the training subjects' respective eyes. The set of training images can include (for example) one or more color fundus photographs, one or more OCT images, and / or one or more other images. Each of the set of labels can include (for example) a visual acuity metric such as a numerical visual acuity, an identification of a visual acuity range, and / or an indication of whether the training subject's eye has a visual acuity below a particular threshold. The visual acuity metric can characterize the visual acuity of the training subject's eye at the time the training images were collected (e.g., on the same day the images were collected) or at a time after the training images were collected (e.g., about 6 months, about 12 months, about 2 years, or about 5 years after the date the training images were collected).

[0094] In some cases, each set of training subjects had been diagnosed with a visual disease or visual medical condition (e.g., prior to collection of the training images), such as macular degeneration, age-related macular degeneration, neovascular macular degeneration, glaucoma, diabetic retinopathy, cataract, or macular edema. In some cases, each of one or more or all of the set of training subjects had not been diagnosed with a visual disease or visual medical condition. In some cases, the set of training subjects includes both training subjects with a visual disease or visual condition and training subjects without a visual disease or visual condition.

[0095] The machine learning model may include a model disclosed herein, such as a deep convolutional neural network. The machine learning model may include one or more residual connections and / or one or more feedforward layers. The set of parameters may include a set of weights and / or one or more convolution kernels.

[0096] The machine learning model may further include one or more pre-processing functions that can pre-process an image and then send the pre-processed image to (for example) a neural network. The machine learning model may further include one or more post-processing functions that can convert an output from (for example) a neural network into another result. For example, a post-processing function may convert a result into a particular visual acuity scale or identify a range within which a numerical neural network output lies.

[0097] In block 110, an image of at least a portion of an eye of a particular subject is accessed. The particular subject may be different from each subject in the set of training subjects and / or may be one of the set of training subjects. The image accessed in block 110 may include a fundus photograph, a color fundus photograph, an OCT image, or another image. The image may include a digital image. The at least a portion of the eye may include (for example) at least a portion of the retina, macula, optic disc, lens, pupil, and / or iris. For example, the image may depict at least a portion of the retina and at least a portion of the macula.

[0098] At block 115, an image of at least a portion of a particular subject's eye is input to a trained machine learning model. The trained machine learning model can determine a visual acuity metric corresponding to a predicted visual acuity of the eye based on the image. The predicted visual acuity can include (for example) a numerical visual acuity, a categorical visual acuity, a visual acuity range, and / or an indication as to whether the predicted visual acuity exceeds a threshold.

[0099] The visual acuity metric is returned at block 120. For example, the visual acuity metric may be presented (e.g., displayed) via a user interface. As another example, the visual acuity metric may be transmitted to another device (e.g., from a server to a user device).

[0100] FIG. 2 illustrates a process 200 for using visual acuity prediction for clinical trials. In block 205, the visual acuity of each of the set of subjects is predicted. The visual acuity may include the visual acuity of the subject's eye. The predicted visual acuity may include a visual acuity metric predicted based on an image of at least a portion of the subject's eye. The visual acuity may be predicted using a method herein, such as some or all of process 100. The predicted visual acuity may be predicted for a time corresponding to the capture of the image or a subsequent time. The predicted visual acuity may be predicted for a specific time corresponding to a clinical trial (e.g., 6 months after the start of treatment, 12 months after the start of monitoring, etc.).

[0101] In block 210, for each of the set of subjects, it can be determined whether the subject is eligible to participate in a particular clinical trial. The determination can be based on whether each of the set of eligibility criteria is met for the subject. A given criterion can include logic such that (for example) it is met if any of multiple conditions are met.

[0102] The set of criteria may indicate that the subject is within a certain age, has been diagnosed with a particular visual disease or condition (e.g., age-related macular degeneration, neovascular age-related macular degeneration), is free of one or more particular other diseases, etc. In some cases, the set of subjects includes subjects enrolled in a clinical trial.

[0103] The set of criteria can include criteria that use the subject's predicted visual acuity. For example, the criteria can indicate that the subject should be associated with a predicted visual acuity metric that exceeds a predetermined threshold. As another example, the criteria can indicate that the subject should be associated with a visual acuity metric change that exceeds a predetermined threshold. The change in visual acuity metric can include an absolute or relative difference between the visual acuity metric associated with a time point and the visual acuity metric associated with a baseline time point. The time point can include (for example) the current time, a predetermined period of time after the start of treatment (e.g., 6 months after the start of treatment, 12 months after the start of treatment), a predetermined period of time after the start of a clinical trial, a specific date, etc. The baseline time point can correspond (for example) to a time before the subject's ocular clinical trial or treatment began, the time when the subject's ocular clinical trial or treatment began, etc. In some cases, the visual acuity metric associated with the time point and the visual acuity metric associated with the time point and the visual acuity metric associated with the baseline time point each include a visual acuity metric generated based on processing at least a portion of an image of the subject's eye (e.g., using a machine learning model). For example, at least a portion of the image of the subject's eye collected at the baseline time point can be used to generate a predicted visual acuity metric for the subject's eye at the baseline time point and also a visual acuity metric for the subject's eye at a subsequent time point. As another example, images of at least a portion of a subject's eye may be collected at each of a baseline time point and a subsequent time point, and the images can be processed to generate an associated predicted visual acuity metric. In some cases, one of the visual acuity metric associated with the time point and the visual acuity metric associated with the time point and the visual acuity metric associated with the baseline time point is generated based on a visual acuity test (e.g., using an eye chart or eye card), and the other of the visual acuity metric associated with the time point and ... baseline time point is generated by processing at least a portion of the images of the subject's eye.

[0104] In block 215, the clinical trial is conducted using a subset of the set of subjects, and each subject in the subset has been determined to be eligible to participate in the clinical trial. Criteria for vision can be determined to be met for each subject in the subset. Conducting the clinical trial can include (for example) stratifying the subset of subjects into two groups and providing the subjects in one group with the investigational treatment. Subjects in the other group can receive (for example) a different treatment, no treatment, a different dosage of the investigational treatment, a different type of administration or formulation of the investigational treatment, standard of care, etc.

[0105] At block 220, results of the clinical trial are generated. For example, the results may indicate the extent to which the investigational treatment was effective in slowing, stopping, or reversing a given medical condition (e.g., as indicated by a visual acuity test, eye imaging, or other test). Efficacy may be estimated by comparing vision or other medical metrics between a first subset and a second subset, between subjects receiving a given treatment and subjects not receiving a given treatment, etc. In some cases, the results indicate the degree of observed visual acuity of subjects who received the treatment over a period of time, compared to the subjects' visual acuity based on processing of images of at least a portion of the subjects' eyes collected at baseline.

[0106] VI. Working Examples VI.A. Machine Learning Models for Processing OCT Images to Predict Visual Acuity Metrics In this example, we evaluated whether deep learning can automatically predict concurrent and future BCVA from optical computed tomography (OCT) images from subjects with neovascular AMD in the Phase 3 HARBOR clinical trial (NCT 00891735, referred to herein as "HARBOR"). Specifically, results were obtained for: (1) a model evaluating the quality of deep learning models to predict accurate best-corrected visual acuity (BCVA) values ​​from OCT images, and (2) a model predicting BCVA of <69 letters (Snellen equivalent, 20 / 40), <59 letters (Snellen equivalent, 20 / 60), or ≤38 letters (Snellen equivalent, 20 / 200) from OCT images. For BCVA outcomes, deep learning models were evaluated for their ability to predict BCVA from OCT images taken at the same (concurrent) visit and for their ability to predict BCVA at 12 months from baseline OCT images. Snellen equivalents of 20 / 40 and 20 / 60 were chosen because visual acuity worse than these levels is considered to reflect visual impairment based on the United States and World Health Organization definitions, respectively. A Snellen equivalent of 20 / 200 or less was used to reflect the definition of legal blindness in the United States.

[0107] VI.A.1. Method VI.A.1.a. Data Sources We used prospectively collected BCVA measurements and OCT images of 1,071 subjects from the Phase 3 HARBOR clinical trial. HARBOR adhered to the tenets of the Declaration of Helsinki and complied with the Health Insurance Portability and Accountability Act. The protocol was approved by each institutional review board before study initiation, and all subjects provided written informed consent for future medical testing and analysis based on the trial results.

[0108] The HARBOR trial enrolled 1,097 adult subjects with treatment-naive subfoveal choroidal neovascularization (CNV) secondary to neovascular AMD, ranging from 20 / 40 to 20 / 320 (Snellen equivalents), using standard ETDRS charts and protocols. Subjects (each in one study eye) were randomized 1:1:1:1 to ranibizumab given according to one of the following treatment regimens: 0.5 mg monthly, 0.5 mg as needed (PRN), 2.0 mg monthly, or 2.0 mg as needed (PRN). Subjects in the PRN group received three monthly injections followed by monthly assessments and underwent retreatment only if there was evidence of disease activity on optical coherence tomography (OCT) imaging or a ≥5-letter decrease in BCVA since the previous visit. BCVA measurements and OCT images were obtained at baseline and at monthly intervals for 24 months.

[0109] OCT images. Images were collected using a spectral-domain Cirrus HD-OCT imaging instrument (Carl Zeiss Meditec, Dublin, CA, USA). The resolution was 200 × 200 × 1024 voxels with dimensions of 30.0 × 30.0 × 2.0 μm, covering a volume of 6 × 6 × 2 mm. The dataset consisted of 50,275 OCT image scans from 1071 subjects. For each of the 50,275 OCT image scans, the retina was flattened to the RPE layer segment provided by the Zeiss software, and the volume was cropped to 384 pixels above and 128 pixels below the flattened RPE. Thirty slices of 512 x 200 pixels were generated per scan by rotating the cropped volume around the z-axis at angles of 0, 30, 60, 90, 120, and 150 degrees relative to the center of the volume and offsetting by -8, -4, 0, 4, and 8 pixels for each of the six angles, resulting in a total of 1,508,250 slices. The OCT image dataset was divided at the subject level into (1) a randomly selected internal validation test set of 147 subjects for use in evaluation (Table 1) and (2) a set of 924 subjects further divided into five for use in model development by cross-validation (Table 2). The subjects in each group remained constant for each outcome variable. [Table 1] [Table 2]

[0110] VI.A.1.b. Deep Learning Modeling Outcome Variables BCVA. The BCVA outcomes of interest were (1) BCVA in ETDRS letters at each study visit and (2) whether specific BCVA values ​​were <69 letters (Snellen equivalent, 20 / 40), <59 letters (Snellen equivalent, 20 / 60), or ≤38 letters (Snellen equivalent, 20 / 200). Snellen equivalents of 20 / 40 and 20 / 60 were chosen because they are thought to reflect functionally significant levels of visual impairment, and Snellen equivalents of 20 / 200 or less were used to reflect the definition of legal blindness in the United States.

[0111] The mean (±standard deviation [SD]) BCVA of the test eye in the internal validation test set was 53.93 (±13.20) letters at baseline and 65.02 (±17.12) letters at 24 months; the mean (±SD) BCVA of the fellow eye was 69.46 (±22.92) letters at baseline and 68.56 (±22.56) letters at 24 months (Table 1). The visual acuity ranges of the test eye at baseline, 6 months, 12 months, 18 months, and 24 months were 55, 83, 89, 82, and 78 letters, respectively (Table 1). The visual acuity ranges of the fellow eye at baseline, 6 months, 12 months, 18 months, and 24 months were 93, 98, 100, 96, and 98 letters, respectively (Table 1).

[0112] VI.A.1.c. Deep Learning Algorithms The DL model was evaluated for its ability to predict (1) accurate BCVA values ​​in ETDRS letters from OCT images obtained at the same visit; (2) accurate BCVA at 12 months from baseline OCT images; (3) BCVA <69 letters (Snellen equivalent, 20 / 40), <59 letters (Snellen equivalent, 20 / 60), or ≤38 letters (Snellen equivalent, 20 / 200) from OCT images obtained at the same visit; and (4) BCVA <69, <59, or ≤38 letters at 12 months from baseline OCT images.

[0113] Prediction of BCVA at co-visits. Deep learning modeling was performed using TensorFlow (1.14.0) with Keras (2.2.5) and ResNet-50 v2 CNN architecture on an Nvidia V100 GPU. Individual slices of 512 × 200 pixels were randomly shuffled from the training set and fed into the CNN using a batch size of 64 images. The model was trained to predict BCVA at the same visit as the OCT image scan (Figure 3). In some cases, a layer of global average pooling and dropout (0.85) was added to the CNN with L2 regularization (0.05) on the final dense layer using a linear activation function. This model is referred to as the "co-visit regression model" in this example. The loss function was mean squared error, and the optimizer was RAdam. The model was trained for only one epoch for each cross-validation fold.

[0114] A simultaneous-visit regression model architecture was used, except that in some cases the final layer used a softmax activation function with a sparse categorical cross-entropy loss function. Models were initialized with weights from the regression model for each fold. Models were trained for two epochs with a non-trainable base model layer using the Adam optimizer, followed by one additional epoch with a trainable base model layer using a stochastic gradient descent (SGD) optimizer.

[0115] Prediction of BCVA from baseline to 12 months. For the regression task of predicting BCVA at 12 months from baseline OCT images, deep learning modeling was performed using Keras (2.2.5) on an Nvidia V100 GPU and TensorFlow (1.14.0) with the ResNet-50 v2 CNN architecture. Individual slices of 512 × 200 pixels were randomly shuffled from the training set and fed into the CNN using a batch size of 64 images. The model was trained to predict BCVA at the same visit as the OCT image scan (Figure 3). In some cases, a layer of global average pooling and dropout (0.995) was added to the CNN with L2 regularization (0.05) on the final dense layer using a linear activation function. For each fold, the model was initialized with weights from a regression model trained to predict BCVA at the same visit. The first three epochs were trained using the SGD optimizer with untrainable base model layers, and then an additional 1000 epochs were trained with trainable base model layers using the SGD optimizer. This model is referred to in this example as the "12-month regression model."

[0116] For classification, a 12-month regression model architecture was used, except that the final layer used a softmax activation function with a sparse categorical cross-entropy loss function. For each fold, the model was initialized with weights from a regression model trained to predict BCVA at 12 months from baseline. The model was trained for 20 epochs with a non-trainable base model layer using the RAdam optimizer. The weights used for prediction were selected from the epoch with the lowest validation loss in each fold.

[0117] Deep Learning Model Evaluation. Metrics assessing model fit at a particular visit were calculated at the eye level by averaging the predictions per eye generated for 30 slices from each of the five development models in the out-of-sample internal validation test set. In other words, each of the five cross-validation models that saw 80% of the data from the training set was used to generate a prediction for each of the 30 slices per eye in the test set, resulting in 150 predictions per eye, which were then averaged. Furthermore, to assess model performance at concurrent visits across all visits while accounting for potential bias from repeated measurements of the same eye, the average of all visits was determined for the regression task, and visits were randomly selected for each subject for the classification task. R from Bland-Altman plots 2 The value, root mean square error (RMSE), and mean difference (MD), as well as the 95% limits of agreement (LOA), were used to evaluate the deep learning regression models, whereas the area under the receiver operating characteristic curve (AUC) and area under the precision-recall curve (AUPRC) were used to evaluate the performance of the deep learning models for classification. To understand whether deep learning prediction of BCVA at 12 months from baseline OCT images contributes additional information compared to using baseline BCVA alone, the R statistical programming language was used to fit linear models to predict BCVA at 12 months from (i) the univariate input of the deep learning prediction of BCVA at 12 months from baseline OCT images, (ii) the univariate input of baseline BCVA, and (iii) the multivariate input of both the deep learning prediction of BCVA at 12 months from baseline OCT images and baseline BCVA. Additionally, results from a 5-fold cross-validation calibration set are reported using the average of 30 adjusted predictions per eye per visit (Tables 3-6).

[0118] VI.A.2. Results VI.A.2.a. Prediction of best corrected visual acuity (BCVA) at concurrent visits Regression results. In the study eye, the deep learning model for predicting BCVA at concurrent visits was significantly higher than the baseline model. 2= 0.24, RMSE = 11.55, MD = -1.81 letters, 95% LOA ranged from -26.57 to 22.95 letters, and for the average across all visits, R 2 = 0.67, RMSE = 8.60, MD = 0.04 letters, and 95% LOA ranged from -16.96 to 17.04 letters (Table 3; Figure 4A). 2 = 0.80, RMSE = 10.35, MD = -1.86 letters, 95% LOA ranged from -23.03 to 19.31 letters, and for the average across all visits, R 2 = 0.84, RMSE = 9.01, MD = 0.51 letters, and 95% LOA ranged from -17.58 to 18.61 letters (Table 3; Figure 4B). 2 = 0.66, RMSE = 11.75, MD = -1.84 letters, 95% LOA ranged from -24.83 to 21.16 letters, averaged across all visits, R 2 = 0.79, RMSE = 8.78, MD = 0.28 letters, and 95% LOA ranged from -17.26 to 17.81 letters (Table 3; Figure 4C). [Table 3] [Table 4] [Table 5] [Table 6]

[0119] Classification model for predicting BCVA <69 letters (Snellen equivalent, 20 / 40). In the study eye, the deep learning model for predicting BCVA <69 letters (Snellen equivalent, 20 / 40) had an AUC = 0.89 and an AUPRC = 0.88 for one randomized concurrent visit per eye, with a class balance of 72 positive and 75 negative eyes (Table 4; Figure 5A). In the fellow eye, the deep learning model for predicting BCVA <69 letters (Snellen equivalent, 20 / 40) had an AUC = 0.93 and an AUPRC = 0.97 for one randomized concurrent visit per eye, with a class balance of 103 positive and 44 negative eyes (Table 4; Figure 5B). In all eyes, the deep learning model for predicting BCVA <69 letters (Snellen equivalent, 20 / 40) had an AUC = 0.92 and an AUPRC = 0.94 for one randomized concurrent visit per eye, with a class balance of 175 positive and 119 negative eyes (Table 4; Figure 5C).

[0120] Classification model for predicting BCVA <59 letters (Snellen equivalent, 20 / 60). In the study eye, the deep learning model for predicting BCVA <59 letters had an AUC = 0.92 and an AUPRC = 0.95 for one randomized concurrent visit per eye, with a class balance of 100 positive and 47 negative eyes (Table 4). In the fellow eye, the deep learning model for predicting BCVA <59 letters had an AUC = 0.97 and an AUPRC = 0.99 for one randomized concurrent visit per eye, with a class balance of 114 positive and 33 negative eyes (Table 4). In all eyes, the deep learning model for predicting BCVA <59 letters had an AUC = 0.95 and an AUPRC = 0.98 for one randomized concurrent visit per eye, with a class balance of 214 positive and 80 negative eyes (Table 4).

[0121] Classification model for predicting BCVA of ≤38 letters (Snellen equivalent, 20 / 200). In the study eye, the deep learning model for predicting BCVA of ≤38 letters had an AUC of 0.92 and an AUPRC of 0.99 for one randomized concurrent visit per eye, with a class balance of 113 positive and 14 negative eyes (Table 4). In the fellow eye, the deep learning model for predicting BCVA of ≤38 letters had an AUC of 0.98 and an AUPRC of 1.00 for one randomized concurrent visit per eye, with a class balance of 129 positive and 18 negative eyes (Table 4). In all eyes, the deep learning model for predicting BCVA of ≤38 letters had an AUC of 0.96 and an AUPRC of 0.99 for one randomized concurrent visit per eye, with a class balance of 262 positive and 32 negative eyes (Table 4).

[0122] VI.A.2.b. Prediction of BCVA at 12 months from baseline OCT images Regression results. The characteristics of the datasets used to evaluate the ability of the models to predict BCVA at 12 months from baseline OCT images are shown in Table 1. The deep learning models for predicting BCVA at 12 months from baseline OCT images performed significantly better than R for the test eye, fellow eye, and all eyes, respectively. 2 = 0.33, 0.75, and 0.58, and RMSE = 14.16, 11.27, and 13.25 (Table 7; Figures 6A, 6B, and 6C). The deep learning model for predicting BCVA at 12 months from baseline OCT images had an MD of -1.63 letters and a 95% LOA of -29.48 to 26.22 letters for the study eye, an MD of -2.31 letters and a 95% LOA of -29.96 to 25.33 letters for the fellow eye, and an MD of -1.97 letters and a 95% LOA of -29.67 to 25.73 letters for all eyes (Figures 4A, 4B, and 4C). The multivariable linear model for predicting BCVA at 12 months from both the deep learning prediction of BCVA at 12 months from baseline OCT images and baseline BCVA had an R of 0.01 for the study eye, fellow eye, and all eyes, respectively. 2= 0.40, 0.88 and 0.68 (Table 7).

[0123] Classification model predicting BCVA <69 letters (Snellen equivalent, 20 / 40). The deep learning model for predicting BCVA at 12 months for <69 letters from baseline OCT images had AUCs of 0.80, 0.92, and 0.87 for the study eye, fellow eye, and all eyes, respectively (Table 8; Figures 7A, 7B, and 7C). In the study eye, the deep learning model predicting BCVA at 12 months for <69 letters from baseline OCT images had an AUPRC of 0.74, with a class balance of 58 positive and 68 negative eyes. In the fellow eye, the deep learning model predicting BCVA at 12 months for <69 letters from baseline OCT images had an AUPRC of 0.97, with a class balance of 88 positive and 37 negative eyes. In all eyes, the deep learning model predicting BCVA at 12 months of <69 letters from baseline OCT images had an AUPRC of 0.91, with a class balance of 146 positive and 105 negative eyes.

[0124] Classification model predicting BCVA of <59 letters (Snellen equivalent, 20 / 60). The deep learning model predicting BCVA at 12 months of <59 letters from baseline OCT images had AUCs of 0.84, 0.93, and 0.89 for the study eye, fellow eye, and all eyes, respectively (Table 8). In the study eye, the deep learning model predicting BCVA at 12 months of <59 letters from baseline OCT images had an AUPRC of 0.90, with a class balance of 83 positive eyes and 43 negative eyes. In the fellow eye, the deep learning model predicting BCVA at 12 months of <59 letters from baseline OCT images had an AUPRC of 0.98, with a class balance of 101 positive eyes and 24 negative eyes. In all eyes, the deep learning model predicting BCVA at 12 months of <59 letters from baseline OCT images had an AUPRC=0.95, with a class balance of 184 positive and 67 negative eyes. [Table 7] [Table 8]

[0125] Classification model predicting BCVA of ≤38 letters (Snellen equivalent, 20 / 200). The deep learning model predicting BCVA at 12 months of ≤38 letters from baseline OCT images had AUCs of 0.77, 0.96, and 0.89 for the study eye, fellow eye, and all eyes, respectively (Table 8). In the study eye, the deep learning model predicting BCVA at 12 months of ≤38 letters from baseline OCT images had an AUPRC of 0.97, with a class balance of 114 positive eyes and 12 negative eyes. In the fellow eye, the deep learning model predicting BCVA at 12 months of ≤38 letters from baseline OCT images had an AUPRC of 0.99, with a class balance of 109 positive eyes and 16 negative eyes. In all eyes, the deep learning model predicting BCVA at 12 months of ≤38 letters from baseline OCT images had an AUPRC = 0.98, with a class balance of 223 positive and 28 negative eyes.

[0126] VI.A.3. Discussion As demonstrated by the above results, the deep learning model is capable of predicting BCVA from OCT images in subjects with neovascular AMD. The predictive accuracy of the derived model was greatest in the fellow eye, reaching a correlation between the mean predicted and mean observed BCVA of ~0.92 (R 2 = 0.84, RMSE = 9.01; Table 3). A moderate to strong correlation was also observed for the BCVA results of the study eye, with a mean of approximately 0.49 (R 2 = 0.24, RMSE = 11.55; Table 3) and at 24 months (post-treatment) approximately 0.79 (R 2 = 0.62, RMSE = 10.54; Table 3).

[0127] To benchmark the presented model results (RMSE 10 letters), we generated predictions using simple linear regression based on screening BCVA versus baseline BCVA in the HARBOR trial. The mean number of days between screening and baseline visits was 8.3 days, with a mean change of 0.6 letters. For this BCVA versus BCVA prediction, the RMSE (from the regression) is 6.1 letters. This mean error of 6.1 letters likely represents an information limit inherent to BCVA observations in the HARBOR trial. Previous studies on the intersession reproducibility of visual acuity scores in neovascular AMD have reported errors of approximately 12 letters. Of note, mean BCVA improvement in response to aVEGF treatment in the HARBOR trial was reported as 7.6 to 9.1 letters. The quality of the deep learning prediction model used in this example, which predicts from OCT images versus BCVA, should be compared to the information limits mentioned above.

[0128] The results suggest the existence of a mapping (defined by deep learning models) between retinal structure and visual function in neovascular AMD. Consequently, OCT images may provide a means of indirectly measuring visual function in clinical trials as well as in evolving clinical practice settings such as telemedicine or home monitoring.

[0129] The difference in model performance between predicting visual function in the test eye and the fellow eye was expected due to the limited range of BCVA at baseline in the test eye. Specifically, HARBOR eligibility criteria required the test eye to have some visual loss and subfoveal CNV at baseline and a BCVA of 20 / 40 to 20 / 320 (Snellen equivalent). These criteria were not required for the fellow eye. Thus, the limited range of BCVA in the test eye at baseline (SD = 13.2; Table 1) reduced the dynamic range and resulted in a more difficult regression task compared to the task of predicting BCVA in the fellow eye, where BCVA variability at baseline was greater (baseline SD = 22.9; Table 1). This observation is supported by the fact that the predictive accuracy of simultaneous prediction of BCVA from OCT images increased in the test eye over the course of the trial, along with increasing BCVA variability (Tables 1-3, Table 9). This increase in range after treatment is consistent with an effective treatment resulting in visual improvement in many subjects. Simulations showed that when the variance of BCVA was (artificially) limited in the fellow eye relative to the variance of the test eye at baseline, i.e., limited to SD = 13.2, R 2 The log(var(BCVA)) was calculated as log(1-R 2 ), the resulting (best fit) pattern was linear (with a slope of -1.22), yielding an R 2 This is consistent with the theory of the near future (Figure 9). Furthermore, the models appear to have similar performance at each time point, as can be seen from the residual error of estimation (RMSE) (Tables 3 and 9). [Table 9]

[0130] Determining the precise relationship between specific and measurable anatomical changes and vision is challenging. The traditional approach to investigating this relationship is to select one or more anatomical features and then perform an analysis to determine whether the association with vision can be quantified. This approach is therefore limited by the researcher's ability to predetermine a potentially large set of retinal structures and features most likely to have a significant relationship with visual function. In contrast, deep learning-based algorithms (especially CNNs) do not require the identification of anatomical features prior to quantitative investigation. Instead, deep learning algorithms evaluate OCT images as a whole and learn directly from the images to identify features that enable the most accurate prediction of the outcome of interest. R is used when using known imaging features along with baseline BCVA to predict BCVA at 12 months. 2 Compared to previously reported cross-validation results on a subset of 614 subjects from the HARBOR trial reporting R = 0.34, the regression deep learning model tested in this example achieved R = 0.34 on the 924 subjects in the training set. 2 = 0.45, R for 126 subjects in the internal validation test set 2 = 0.40. In principle, this could lead to the identification of previously unrecognized anatomical features, or combinations of features, that are important for visual function.

[0131] A separate deep learning model was able to predict BCVA values ​​in the study eye at 12 months from the time of baseline OCT image measurement, with a moderate correlation of approximately 0.57 (R 2 = 0.33; Table 7). It is interesting to note that when added to a regression model that already included baseline BCVA (P < 0.001), the OCT image-based prediction (from baseline) remained highly statistically significant (P < 0.001). In this multivariable model, both predictors contributed significantly to the model R 2= 0.40 (Table 7), providing approximately equal information on future visual function. When used as a stratification factor at baseline, this predictive model can be translated to smaller / shorter trials with the same statistical power. Initially, separate models were used for the test eye and fellow eye, but surprisingly, results were not significantly different in terms of predictive performance from models trained on the test eye and fellow eye pooled together.

[0132] The use of deep learning models to predict BCVA could have meaningful clinical utility. Measuring BCVA is often cumbersome and requires specialized resources for accurate refractive testing. Indeed, for assessing retinal health outside of a clinic setting, the ability to augment visual function measurements with computer vision-based analysis of OCT images is likely to prove beneficial for subject screening and monitoring. For example, remote consultations via telemedicine could be supported, and physicians could use deep learning data about a patient's current and future visual potential to aid clinical decisions. Furthermore, in clinical trials, deep learning models that help predict future BCVA response could be used to aid trial enrollment or trial stratification by focusing on individuals more likely to benefit from treatment.

[0133] VI.B. Machine Learning Models for Processing Color Fundus Images to Predict Visual Acuity Metrics In this example, we evaluated whether deep learning can automatically predict BCVA from color fundus photographs (CFP) images of subjects with neovascular AMD. Specifically, a first deep learning regression model (including a deep convolutional neural network and a linear activation function) was used to accurately predict BCVA from CFP images at chart distances of 2 meters (m) and 4 meters (m). A second deep learning classification model (including a deep convolutional neural network and a softmax activation function) was then used.

[0134] VI.B.1. Methods and Data VI.B.1.a. Data Sources Prospectively collected BCVA measurements and CFP images from 707 subjects from the Phase 3 MARINA clinical trial (NCT 00056836) and 413 subjects from the Phase 3 ANCHOR clinical trial (NCT 00061594) were used. MARINA and ANCHOR were in accordance with the Declaration of Helsinki and Health Insurance Portability and Accountability Act. Protocols were approved by each institutional review board before study initiation, and all subjects provided written informed consent for future medical testing and analysis based on the trial results.

[0135] MARINA enrolled 720 adult subjects with subfoveal choroidal neovascularization (CNV) secondary to neovascular AMD if they had a BCVA of 20 / 40 to 20 / 320 (Snellen equivalent) in the study eye using standard ETDRS charts and protocols. ANCHOR enrolled 426 adult subjects with subfoveal choroidal neovascularization (CNV) secondary to neovascular AMD if they had a BCVA of 20 / 40 to 20 / 320 (Snellen equivalent) in the study eye.

[0136] VI.B.1.b.CFP images A total of 36,541 images from MARINA and 33,591 images from ANCHOR were analyzed. The F1M, F2, and F3M inner capture fields from both left and right stereopsis were included in the analysis, but the outer field of view of the eye (FR capture field) was excluded. (In particular, the F4, F5, F6, and / or F7 fields may alternatively be used to predict visual acuity metrics.) To remove irrelevant information, images were cropped to fit the circle generated by the camera lens. Images were then resized to 299 × 299 × 3 pixels. CFP images from MARINA were divided into five folds at the subject level and used for model development by cross-validation. The subjects in each fold remained constant for both the regression and classification tasks.

[0137] VI.B.1.c. Deep learning modeling outcome variables The BCVA outcomes of interest were (1) BCVA in ETDRS letters at each visit and (2) whether a specific BCVA value was <69 letters (20 / 40 Snellen equivalent). The 20 / 40 Snellen equivalent was chosen because it is thought to reflect a functionally significant level of visual impairment. The mean (±standard deviation [SD]) BCVA of the eyes in the ANCHOR external validation test set was 55.7 ± 24.7 letters for subjects who had BCVA measurements at a distance of 2 m and 55.0 ± 25.2 letters for subjects who had BCVA measurements at a distance of 4 m (Table 10). [Table 10]

[0138] VI.B.1.d. Deep learning algorithms Deep learning models were evaluated for their ability to predict (1) accurate BCVA values ​​for ETDRS letters from CFP images obtained at the same visit and (2) BCVA for <69 letters (Snellen equivalent, 20 / 40) from CFP images obtained at the same visit.

[0139] Deep learning modeling was performed on an Nvidia V100 GPU using TensorFlow (1.14.0) with Keras (2.2.5) and an Inception-ResNet-v2 CNN architecture trained to predict BCVA at the same visit as CFP.

[0140] For the regression model, global average pooling, dropout (0.5), dense (256), and dense (1) layers were added to the base CNN model. To account for the distance at which BCVA was measured, a corresponding chart distance of 2 m or 4 m was concatenated into the final dense layer (Figure 10). The loss function was mean squared error. For each of the five cross-validation folds, the model was initialized with weights pre-trained on the ImageNet dataset and trained for two epochs with the non-trainable base model layers using the Adam optimizer, then for an additional 200 epochs with the trainable base model layers using the RAdam optimizer.

[0141] For the classification model, the architecture remained the same as for the regression model, except that the final layer was dense(2) using a softmax activation function with a cross-entropy loss function for sparse categories. The model was initialized with weights from the regression model for each fold. The model was trained for 3 epochs with untrainable base model layers using the Adam optimizer.

[0142] VI.B.1.e. Evaluate deep learning models. The model weights from the epoch with the lowest validation loss were selected from each cross-validation fold. Metrics for assessing model fit are calculated by averaging the predictions generated for the eye at each visit from each of the five cross-validation folds of the ANCHOR out-of-sample external validation test set (Table 11). Additionally, results are reported for the MARINA five-fold cross-validation training set (Table 11). 2 The value was used to benchmark the deep learning regression models, while the area under the receiver operating characteristic curve (AUC) was used to evaluate the performance of the deep learning models for classification.

[0143] VI.B.2. Results Results of the regression model. The regression model predicting BCVA at a chart distance of 2 m was R for the MARINA deployment set. 2The regression model predicting BCVA at a chart distance of 4 m was R = 0.56 (95% CI: 0.54, 0.57) for the MARINA deployment set, and R = 0.59 (95% CI: 0.57, 0.60) for the ANCHOR external validation test set (Table 11, Figure 11A). 2 = 0.57 (95% CI: 0.55, 0.60), and R 2 = 0.60 (95% CI: 0.57, 0.63) (Table 11, Figure 11B). Figures 11A and 11B show the performance of the deep learning regression model analyzing color fundus photograph images to predict BCVA in the ANCHOR external validation test set. Figure 11A shows the actual vs. predicted BCVA at a chart distance of 2 meters. R 2 =0.59. Figure 11B shows the actual vs. predicted BCVA at a chart distance of 4 meters. R 2 =0.60.

[0144] Classification model results. The classification model predicting BCVA <69 letters (20 / 40 Snellen equivalent) at a chart distance of 2 m had an AUC of 0.86 (95% CI: 0.85, 0.87) for the MARINA development set and an AUC of 0.86 (95% CI: 0.85, 0.87) for the ANCHOR external validation test set (Table 11, Figure 12A). The classification model predicting BCVA <69 letters (20 / 40 Snellen equivalent) at a chart distance of 4 m had an AUC of 0.87 (95% CI: 0.85, 0.88) for the MARINA development set and an AUC of 0.88 (95% CI: 0.86, 0.90) for the ANCHOR external validation test set (Table 11, Figure 12B). Figures 12A and 12B show the performance of a deep learning classification model analyzing color fundus photographs to predict BCVA <69 letters (Snellen equivalent <20 / 40) on the ANCHOR external validation test set. Figure 12A shows the area under the receiver operating characteristic curve (AUC) at a chart distance of 2 meters. AUC=0.86. Figure 12B shows the AUC at a chart distance of 4 meters. AUC=0.88. [Table 11]

[0145] VI.B.3. Discussion The results demonstrate that neural networks are capable of learning quantitative relationships between retinal structure and visual function in subjects with neovascular age-related macular degeneration.

[0146] VII. Computing Systems FIG. 13 illustrates a network 1300 of computing systems that can be configured to perform some or all of one or more operations and / or one or more methods disclosed herein. The network 1300 can include one or more ocular imaging systems 1305 configured to collect one or more images of a subject's eye. The one or more ocular imaging systems can include one or more techniques disclosed herein (e.g., optical coherence tomography as disclosed in Section II.A.1 or color fundus photography as disclosed in Section II.A.2). The ocular imaging system can include optical components configured to collect (for example) OCT images or color fundus photography. For example, the ocular imaging system can include an interferometer (e.g., a Michelson-type interferometer), a light source (e.g., low coherence, broad bandwidth), and a beam splitter. As another example, the ocular imaging system can include a fundus camera. The ocular imaging system can include a computing system (e.g., having one or more processors, one or more memories, and / or one or more transmitters) for storing and / or transmitting the image (or a processed version thereof).

[0147] Network 1300 may include one or more visual function assessment systems 1310, which may include one or more computing systems (e.g., having one or more processors, one or more memories, or one or more transmitters) configured to collect and transmit visual acuity metrics generated based on a subject's responses to viewing visual stimuli, such as an eye chart or eye card. The visual acuity metrics may include the metrics disclosed in Section II.C.1. The visual acuity metrics may be determined based on the techniques disclosed herein. In some cases, the visual function assessment system presents visual stimuli on a screen of the system. In some cases, the visual stimuli are presented separately (e.g., via a physical chart or card).

[0148] The machine learning model system 1315 may include one or more computing systems (e.g., having one or more processors, one or more memories, one or more transmitters, and / or one or more receivers). The machine learning model system 1315 may be partially or entirely a cloud computing system. The machine learning model system 1315 may include one or more servers. In some cases, the machine learning model system 1315 is or is included within a subject device and / or medical device (e.g., a wearable device, a smartphone, etc.) of the subject.

[0149] The machine learning model system 1315 can be configured to train and / or use a machine learning model, which can include one or more preprocessing functions (e.g., as disclosed in Section II.B.) and one or more neural networks (e.g., having the architecture and / or characteristics disclosed in Section III.A.).

[0150] The machine learning model system 1315 can train the machine learning model using (for example) the techniques disclosed in Section III.B. The machine learning model system 1315 can receive training data (for example, as disclosed in Section II.C) for each of a set of training subjects. The training data can include, for each training subject, one or more images of one or both eyes (from the eye imaging system 1305) and a corresponding visual acuity score or metric for one or both eyes (from the visual function assessment system 1310). In some cases, multiple visual acuity metrics are received for each of the one or both eyes, corresponding to different time points (e.g., relative to the time the eye images were collected). For example, one metric can correspond to a visual function test (e.g., eye chart reading) administered on the same day the eye images were collected, and another metric can correspond to a visual function test administered about six months or about one year after the eye images were collected. The machine learning model system 1315 can use the images and the visual acuity metric to train a machine-learning model (e.g., using a deep convolutional network) to predict the visual acuity metric (or another type of visual acuity metric, such as a binary indicator) based on the eye images. In some cases, the model includes one or more preprocessing functions to preprocess the images. Training the model can include learning a set of parameters.

[0151] The machine learning model system 1315 can then receive another image (e.g., from another eye imaging system 1305) corresponding to another subject, and can use the trained model to predict the subject's current and / or future vision metrics associated with the other image.

[0152] The machine learning model system 1315 can transmit the predicted visual acuity metrics (e.g., along with the subject's identifier) ​​to one or more user devices 1320. The one or more user devices 1320 may correspond to the subject or the subject's caregiver (for example). Users of the one or more user devices 1320 can use the results in the manner disclosed in Section IV (for example).

[0153] In some cases, two or more of the eye imaging system 1305, visual function assessment system 1310, or user device 1320 are owned by the same entity, co-located, and / or included within the same computing system (e.g., in a care provider's office). In some cases, the subject's device (e.g., a smartphone) includes at least a portion of each of two or more of the illustrated components. For example, an attachment to the smartphone can be used to collect images of the subject's eyes, a trained machine learning model can be run locally on the smartphone, and predicted visual acuity can be provided to the smartphone.

[0154] VIII. Illustrative Embodiments A first exemplary embodiment is a method for predicting visual acuity based on image processing. The method includes: accessing images of at least a portion of a subject's eye; inputting the images into a machine learning model to determine a predicted visual acuity metric corresponding to the predicted visual acuity, the machine learning model including a set of parameters determined using a set of training images, each of the set of training images depicting at least a portion of an eye of a training subject in the set of training subjects, and a set of labels identifying observed visual acuity for each of the set of training subjects; and a function that associates the images and parameters with the visual acuity metric; further using the function that associates the images and parameters with the visual acuity metric (e.g., the function is learned to associate training images with observed visual acuity metrics, and the function can be used to associate non-training images with predicted visual acuity metrics); and returning the predicted visual acuity metric.

[0155] A second exemplary embodiment includes the first exemplary embodiment, wherein the image is an optical coherence tomography (OCT) image.

[0156] A third exemplary embodiment includes the first exemplary embodiment, where the image is a color fundus photograph image.

[0157] A fourth example embodiment includes any of the first through third example embodiments, wherein the predicted visual acuity metric corresponds to the predicted visual acuity on a baseline date when the image was captured by the imaging device.

[0158] A fifth exemplary embodiment includes any of the first through third exemplary embodiments, wherein the predicted visual acuity metric includes one or more numbers, the one or more numbers representing predicted visual acuity for a date at least six months from the baseline date when the image was captured.

[0159] A sixth illustrative example includes any of the first through third illustrative examples, wherein the predicted visual acuity metric is a binary value representing a prediction as to whether the subject's visual acuity on the baseline day when the image was captured by the imaging device is worse than a threshold visual acuity value.

[0160] A seventh exemplary embodiment includes any of the first through third exemplary embodiments, wherein the predicted visual acuity metric is a binary value representing a prediction as to whether the subject's visual acuity at least six months from the baseline date when the image was captured by the imaging device will be worse than a threshold visual acuity value.

[0161] An eighth exemplary embodiment includes the sixth or seventh exemplary embodiment, wherein the threshold corresponds to a Snellen fraction of 20 / 160, 20 / 80, or 20 / 40.

[0162] A ninth illustrative example provides a method for, in response to the predicted visual acuity metric,

[0163] Any of the first through eighth exemplary embodiments further comprising providing a recommendation that the subject receive a pharmacological treatment.

[0164] A tenth exemplary embodiment is a method for detecting a plurality of images, the method comprising:

[0165] further comprising slicing the image into a plurality of two-dimensional slices, each of the slices being acquired at a different scan central angle and offset a different number of pixels from each other;

[0166] Any of the first through ninth exemplary embodiments, wherein inputting the image into the model includes inputting a slice into the model.

[0167] An eleventh exemplary embodiment includes any of the first through tenth exemplary embodiments, wherein the model is a deep learning model.

[0168] A twelfth exemplary embodiment includes any of the first through eleventh exemplary embodiments, wherein the model is or includes a convolutional neural network.

[0169] A thirteenth exemplary embodiment includes any of the first through twelfth exemplary embodiments, wherein the model uses a set of convolution kernels.

[0170] A fourteenth exemplary embodiment includes any of the first through thirteenth exemplary embodiments, wherein the model includes a ResNet model.

[0171] A fifteenth exemplary embodiment includes any of the first through fourteenth exemplary embodiments, wherein the model is an Inception model.

[0172] A sixteenth exemplary embodiment includes any of the first through fifteenth exemplary embodiments, wherein the predicted visual acuity metric corresponds to a predicted visual acuity of the subject's eye.

[0173] A seventeenth exemplary embodiment includes any of the first through sixteenth exemplary embodiments, wherein the predicted visual acuity metric corresponds to a predicted corrected visual acuity that predicts the subject's visual acuity while wearing glasses or contacts.

[0174] An eighteenth exemplary embodiment includes any of the first through seventeenth exemplary embodiments, wherein the predicted visual acuity metric corresponds to a predicted best-corrected visual acuity of the subject's eye.

[0175] A nineteenth exemplary embodiment includes any of the first through seventeenth exemplary embodiments, wherein the subject was previously diagnosed with age-related macular degeneration at the time the images were collected.

[0176] A twentieth exemplary embodiment includes any of the first through seventeenth exemplary embodiments, wherein the subject was previously diagnosed with neovascular age-related macular degeneration at the time the images were collected.

[0177] A twenty-first exemplary embodiment includes any of the first through seventeenth exemplary embodiments, wherein the subject was previously diagnosed with dry age-related macular degeneration at the time the images were collected.

[0178] A twenty-second exemplary embodiment includes any of the first through seventeenth exemplary embodiments, wherein each of the training subjects was previously diagnosed with age-related macular degeneration before the training images of the set of images were collected.

[0179] A twenty-third exemplary embodiment includes any of the first through seventeenth exemplary embodiments, wherein each of the training subjects was previously diagnosed with neovascular age-related macular degeneration before the training images of the set of images were collected.

[0180] A twenty-fourth exemplary embodiment includes any of the first through seventeenth exemplary embodiments, wherein each of the training subjects was previously diagnosed with dry age-related macular degeneration before the training images of the set of images were collected.

[0181] A twenty-fifth exemplary embodiment is

[0182] Any of the first through twenty-fourth exemplary embodiments further includes training a model using the set of training images and the set of labels.

[0183] A twenty-sixth exemplary embodiment includes any of the first through twenty-fifth exemplary embodiments, wherein the machine learning model includes one or more preprocessing functions and one or more neural networks.

[0184] A twenty-seventh exemplary embodiment includes any of the first through twenty-fifth exemplary embodiments, wherein the image of at least a portion of the eye includes a preprocessed version of the eye generated by applying one or more preprocessing functions to a raw image of at least a portion of the subject's eye.

[0185] A twenty-eighth exemplary embodiment includes any of the twenty-sixth or twenty-seventh exemplary embodiments, wherein the one or more pre-processing functions include a function that flattens the image.

[0186] A twenty-ninth exemplary embodiment includes any of the twenty-sixth or twenty-seventh exemplary embodiments, wherein the one or more pre-processing functions include a function that generates one or more B-scan images or one or more C-scan images.

[0187] A thirtieth exemplary embodiment includes either the twenty-sixth or twenty-ninth exemplary embodiment, wherein the one or more pre-processing functions include a trimming function.

[0188] A thirty-first exemplary embodiment is a method for predicting a visual acuity of an eye of a subject at a first distance, the method comprising:

[0189] determining another predicted visual acuity metric corresponding to the predicted visual acuity of the same or a different eye of the subject using the machine learning model;

[0190] and returning another predicted visual acuity metric.

[0191] A thirty-second exemplary embodiment is a method of stratifying a clinical trial, comprising: performing, for each subject of a set of subjects, the method of predicting visual acuity based on image processing of any of the first through thirty-first exemplary embodiments; determining, for each of the set of subjects, whether the subject is eligible to participate in the clinical trial, wherein the determination is based on whether each of a set of eligibility criteria is met for the subject, and wherein evaluation of a particular eligibility criterion of the set of eligibility criteria is made using the subject's predicted visual acuity; and conducting the clinical trial with a subset of the set of subjects, wherein each subject in the subset is determined to be eligible to participate in the clinical trial.

[0192] A thirty-third exemplary embodiment includes the thirty-second exemplary embodiment, further including conducting the clinical trial according to the assignment of subjects to the first subject group and the second subject group.

[0193] A thirty-fourth exemplary embodiment includes selecting a treatment from a set of potential treatments for a subject based on a predicted visual acuity metric determined according to the method of any of the first through thirty-first exemplary embodiments, and outputting identification information of the treatment.

[0194] A thirty-fifth exemplary embodiment includes selecting a treatment from among a set of potential treatments for a subject based on a predicted visual acuity metric determined according to any of the first through thirty-first exemplary embodiments, and treating the subject using the treatment.

[0195] A thirty-sixth exemplary embodiment includes any of the thirty-fourth or thirty-fifth exemplary embodiments, in which at least some of the training subjects each received treatment before participating in a vision test from which a label corresponding to the training subject was determined.

[0196] A thirty-seventh exemplary embodiment includes any of the thirty-fourth or thirty-fifth exemplary embodiments, in which each of the training subjects received treatment before participating in the vision test from which the label corresponding to the training subject was determined.

[0197] A thirty-eighth exemplary embodiment includes any of the thirty-fourth or thirty-fifth exemplary embodiments, in which at least some of the training subjects each received another treatment different from the treatment prior to participating in the vision test from which the label corresponding to the training subject was determined, and the subjects previously received the other treatment.

[0198] A thirty-ninth exemplary embodiment includes any of the thirty-fourth or thirty-fifth exemplary embodiments, in which each of the training subjects received another treatment different from the treatment prior to participating in the vision test from which the label corresponding to the training subject was determined, and the subject previously received the other treatment.

[0199] A fortieth exemplary embodiment includes one or more data processors and a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of one or more methods disclosed herein.

[0200] A forty-first exemplary embodiment includes a computer program product tangibly embodied in a non-transitory machine-readable storage medium, comprising instructions configured to cause one or more data processors to perform some or all of one or more of the methods disclosed herein.

[0201] IX. Further Considerations Some disclosures herein refer to a subject (e.g., identifying a subject's visual acuity, predicting a subject's visual acuity, treating a subject, etc.), and it will be understood that in some embodiments, some or all of these disclosures relate to a particular eye of a subject.

[0202] Some embodiments of the present disclosure include a system including one or more data processors. In some embodiments, the system includes a non-transitory computer-readable storage medium including instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of one or more methods and / or some or all of one or more processes disclosed herein. Some embodiments of the present disclosure include a computer program product tangibly embodied in a non-transitory machine-readable storage medium including instructions configured to cause one or more data processors to perform some or all of one or more methods and / or some or all of one or more processes disclosed herein.

[0203] The terms and expressions which have been employed are used as terms of description rather than of limitation, and there is no intention in the use of such terms and expressions to exclude equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the invention as claimed. Thus, although the claimed invention has been specifically disclosed by embodiments and optional features, it will be understood that modifications and variations of the concepts disclosed herein may be resorted to by those skilled in the art, and that such modifications and variations are deemed to be within the scope of the invention as defined by the appended claims.

[0204] The description provides only preferred exemplary embodiments and is not intended to limit the scope, applicability, or configuration of the present disclosure. Rather, the description of preferred exemplary embodiments provides those skilled in the art with an enabling description for implementing various embodiments. It will be understood that various changes can be made in the function and arrangement of elements without departing from the spirit and scope of the appended claims.

[0205] Specific details are given in the following description to provide a thorough understanding of the embodiments. However, it will be understood that the embodiments can be practiced without these specific details. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form in order to avoid obscuring the embodiments in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail in order to avoid obscuring the embodiments.

Claims

1. 1. A computer-implemented method for predicting visual acuity based on image processing, comprising: accessing an image of at least a portion of the subject's eye, the image captured by an imaging device on a first date; Inputting the image into a machine learning model, a first predicted visual acuity metric, the subject's visual acuity as determined by administering a visual acuity evaluation to the subject on the first date; and a second predicted visual acuity metric, which is the subject's visual acuity determined by administering the visual acuity evaluation to the subject on a second date that is a specified interval after the first date; and determining the machine learning model: A set of parameters, a set of training images corresponding to a set of training subjects, each image in the set of training images depicting at least a portion of an eye of a corresponding training subject in the set of training subjects; a set of parameters learned using a set of labels identifying the observed visual acuity for each subject in the set of training subjects; a function relating the image and the set of parameters to the first predicted visual acuity metric and the second predicted visual acuity metric; returning the first predicted visual acuity metric and the second predicted visual acuity metric; A method comprising:

2. The computer-implemented method of claim 1 , wherein the image is an optical coherence tomography (OCT) image.

3. The computer-implemented method of claim 1 , wherein the image is a color fundus photograph image.

4. The computer-implemented method of any one of claims 1 to 3, wherein the second date is 12 months after the first date.

5. 4. The computer-implemented method of claim 1, wherein the second predicted visual acuity metric includes one or more numbers, the one or more numbers representing the visual acuity on the second date, the second date being at least six months after the first date on which the image was captured.

6. 4. The computer-implemented method of claim 1, wherein the first predictive visual acuity metric is a binary value that predicts whether the subject's visual acuity on the first date is worse than a threshold visual acuity value.

7. 4. The computer-implemented method of claim 1, wherein the second predictive visual acuity metric is a binary value that predicts whether the subject's visual acuity on the second date at least six months after the first date will be worse than a threshold visual acuity value.

8. 8. The computer-implemented method of claim 6 or 7, wherein the threshold visual acuity value is a Snellen fraction of 20 / 160, 20 / 80, or 20 / 40.

9. in response to the first predicted visual acuity metric and the second predicted visual acuity metric; The computer-implemented method of any one of claims 1 to 8, further comprising providing a recommendation that the subject undergo pharmacological treatment.

10. the image is a three-dimensional image, and the method comprises: further comprising slicing the image into a plurality of two-dimensional slices, each slice of the plurality of two-dimensional slices being acquired at a different scan central angle and offset by a different number of pixels from each other slice; 10. The computer-implemented method of claim 1, wherein inputting the image into the machine learning model comprises inputting the plurality of two-dimensional slices into the machine learning model.

11. 11. The computer-implemented method of claim 1, wherein the machine learning model is a deep learning model.

12. 12. The computer-implemented method of any one of claims 1 to 11, wherein the machine learning model is or comprises a convolutional neural network.

13. 13. The computer-implemented method of any one of claims 1 to 12, wherein the machine learning model uses a set of convolution kernels.

14. 14. The computer-implemented method of claim 1, wherein the machine learning model comprises at least one of a ResNet model or an Inception model.

15. 15. The computer-implemented method of any one of claims 1 to 14, wherein the visual acuity assessment evaluates the subject's response to one or more visual stimuli.

16. 16. The computer-implemented method of any one of claims 1 to 15, wherein the image of the subject's eye is captured when the subject is not wearing glasses or contacts.

17. 17. The computer-implemented method of any one of claims 1 to 16, wherein the image of the subject's eye is captured while the subject is wearing glasses or contacts.

18. 18. The computer-implemented method of any one of claims 1 to 17, wherein each of the first predicted visual acuity metric and the second predicted visual acuity metric comprises a predicted best-corrected visual acuity of the eye of the subject.

19. 18. The computer-implemented method of any one of claims 1 to 17, wherein the subject has been previously diagnosed with age-related macular degeneration at the time the images are collected.

20. 18. The computer-implemented method of any one of claims 1 to 17, wherein the subject has been previously diagnosed with neovascular age-related macular degeneration at the time the image is collected.

21. 18. The computer-implemented method of any one of claims 1 to 17, wherein the subject has been previously diagnosed with dry age-related macular degeneration at the time the images are collected.

22. 18. The computer-implemented method of claim 1, wherein each of the training subjects is a subject who was previously diagnosed with age-related macular degeneration before the training images of the set of images were collected.

23. 18. The computer-implemented method of any one of claims 1 to 17, wherein each of the training subjects is a subject who was previously diagnosed with neovascular age-related macular degeneration before the training images of the set of images were collected.

24. 18. The computer-implemented method of any one of claims 1 to 17, wherein each of the training subjects is a subject who was previously diagnosed with dry age-related macular degeneration before the training images of the set of images were collected.

25. 25. The computer-implemented method of any one of claims 1 to 24, further comprising training the machine learning model using the set of training images and the set of labels.

26. 26. The computer-implemented method of any one of claims 1 to 25, wherein the machine learning model comprises one or more pre-processing functions and one or more neural networks.

27. 26. The computer-implemented method of any one of claims 1 to 25, wherein the image of the at least part of the eye comprises a preprocessed version of the eye generated by applying one or more preprocessing functions to a raw image of the at least part of the eye of the subject.

28. 28. The computer-implemented method of claim 26 or 27, wherein the one or more pre-processing functions include a function that flattens the image.

29. 29. The computer implemented method of any one of claims 26 to 28, wherein the one or more pre-processing functions include a function that generates one or more B-scan images or one or more C-scan images.

30. 30. The computer-implemented method of any one of claims 26 to 29, wherein the one or more pre-processing functions comprises a trimming function.

31. the first predicted visual acuity metric and the second predicted visual acuity metric correspond to visual acuity of the eye of the subject at a first distance, and the method further comprises: Using the machine learning model to determine a third predicted visual acuity metric, the subject's visual acuity as determined by administering the visual acuity evaluation to the subject either (1) on a third date different from the first and second dates for the same eye of the subject, or (2) on a different eye of the subject; returning the third predicted visual acuity metric; and 30. The computer-implemented method of any one of claims 26 to 29, further comprising:

32. 1. A computer-implemented method for stratifying a clinical trial, comprising: - for each subject of a set of subjects, performing the method for visual acuity prediction based on image processing according to any one of claims 1 to 31; determining, for each of the set of subjects, whether the subject is eligible to participate in the clinical trial, wherein the determination is based on whether each of a set of eligibility criteria is met for the subject, and wherein evaluation of a particular eligibility criterion of the set of eligibility criteria is made using at least one of the first predicted visual acuity metric and the second predicted visual acuity metric for the subject; and determining that the clinical trial will be conducted using a subset of the set of subjects, wherein each subject in the subset is a subject determined to be eligible to participate in the clinical trial. method.

33. 33. The computer-implemented method of claim 32, wherein the clinical trial is determined to be conducted according to assignment of subjects to a first subject group and a second subject group.

34. selecting a treatment from a set of potential treatments for the subject based on at least one of the first predicted visual acuity metric or the second predicted visual acuity metric determined according to the method of any one of claims 1 to 31; outputting the identification information of the treatment; A computer-implemented method comprising:

35. 1. A computer-implemented method comprising: selecting a treatment from a set of potential treatments for the subject based on at least one of the first predicted visual acuity metric or the second predicted visual acuity metric determined according to the method of any one of claims 1 to 31, and determining that the subject will be treated using the treatment; A method comprising:

36. 36. The computer-implemented method of claim 34 or 35, wherein at least some of the training subjects are subjects who received the treatment before participating in a vision test from which the set of labels corresponding to the training subject was determined.

37. 36. The computer-implemented method of claim 34 or 35, wherein each of the training subjects is a subject who received the treatment before participating in a vision test from which the set of labels corresponding to the training subject was determined.

38. 36. The computer-implemented method of claim 34 or 35, wherein at least some of the training subjects each received a different treatment than the treatment before participating in a vision test from which the set of labels corresponding to the training subject was determined, and the subjects previously received the different treatment.

39. 36. The computer-implemented method of claim 34 or 35, wherein each of the training subjects is a subject who received another treatment different from the treatment before participating in the vision test from which the set of labels corresponding to the training subject was determined, and the subject is a subject who previously received the other treatment.

40. one or more data processors; a non-transitory computer-readable storage medium comprising instructions that, when executed on said one or more data processors, cause said one or more data processors to perform part or all of one or more methods of any one of claims 1 to 39; A system comprising:

41. comprising instructions configured to cause one or more data processors to carry out part or all of one or more of the methods of any one of claims 1 to 39, A computer program product tangibly embodied in a non-transitory machine-readable storage medium.

Citation Information

Patent Citations

  • System and method for classifying and quantifying age-related macular degeneration

    US20170119243A1

  • Deep learning-based diagnosis and referral of ophthalmic diseases and disorders

    US20190110753A1