Multimodal prediction of visual acuity response
A neural network system processing 2D and 3D imaging data predicts visual acuity response to anti-VEGF therapies for AMD, addressing subject-specific variability and enhancing treatment customization and clinical trial efficiency.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- F HOFFMANN LA ROCHE & CO AG
- Filing Date
- 2021-12-02
- Publication Date
- 2026-04-27
AI Technical Summary
Current anti-VEGF therapies for neovascular AMD have subject-specific responses, necessitating a need for systems and methods to predict visual acuity response accurately to tailor treatment regimens and reduce complications.
Utilizing neural networks to process 2D and 3D imaging data, such as color fundus and OCT images, to predict visual acuity response (VAR) through a multimodal approach, enabling individualized treatment dosing and clinical trial enhancements.
The multimodal approach provides accurate and efficient predictions of visual acuity changes post-treatment, facilitating individualized treatment plans and improving clinical trial design.
Smart Images

Figure 0007851930000005 
Figure 0007851930000006 
Figure 0007851930000007
Abstract
Description
Technical Field
[0001] Cross-reference This application claims priority to U.S. Provisional Patent Application No. 63 / 121,213, filed December 3, 2020, with the title "MULTIMODAL PREDICTION OF VISUAL ACUITY RESPONSE" and U.S. Provisional Patent Application No. 63 / 175,544, filed April 15, 2021, with the title "MULTIMODAL PREDICTION OF VISUAL ACUITY RESPONSE", the entire disclosures of which are incorporated herein by reference for all purposes.
[0002] Field This description generally relates to predicting visual acuity response in subjects diagnosed with age-related macular degeneration (AMD). More specifically, this description provides methods and systems for predicting visual acuity response in subjects diagnosed with AMD using information obtained from one or more imaging modalities.
Background Art
[0003] Introduction Age-related macular degeneration (AMD) is a disease that affects the central region of the retina of the eye called the macula. AMD is the leading cause of vision loss in subjects over 50 years of age. Neovascular AMD (nAMD) is one of the two progressive stages of AMD. In nAMD, new and abnormal blood vessels proliferate uncontrollably under the macula. This type of proliferation can cause swelling, hemorrhage, fibrosis, other problems, or a combination thereof. Treatment for nAMD typically involves anti-vascular endothelial growth factor (anti-VEGF) therapy (e.g., anti-VEGF drugs such as ranibizumab). The retinal response to such treatment is at least partially subject-specific, and as a result, different subjects may respond differently to the same type of anti-VEGF drug. Furthermore, anti-VEGF therapy is typically administered by intravitreal injection, which is expensive and can itself cause complications (e.g., blindness). Therefore, there is a need for systems and methods that can predict how well subjects with nAMD are likely to respond to treatment with anti-VEGF drugs. [Overview of the Initiative]
[0004] overview This disclosure provides systems and methods for predicting visual acuity response (VAR). The systems and methods generally utilize neural networks. In some embodiments, the systems and methods utilize a neural network configured to receive inputs including two-dimensional (2D) imaging data, such as color fundus imaging (CFI) data, and apply a trained model to the inputs to predict VAR responses (e.g., predicted changes in a subject's visual acuity in response to a treatment, such as treatment with an anti-VEGF drug). In some embodiments, the systems and methods utilize a neural network configured to receive inputs including three-dimensional (3D) imaging data, such as optical coherence tomography (OCT) data, and apply a trained model to the inputs to predict VAR responses. In some embodiments, the methods and systems are configured to receive a first input including 2D imaging data and a second input including 3D imaging data, and to apply a trained model to the first and second inputs to predict VAR responses. [Brief explanation of the drawing]
[0005] For a more complete understanding of the principles and advantages disclosed herein, refer to the following description in conjunction with the accompanying drawings.
[0006] [Figure 1] This is a block diagram of a prediction system according to various embodiments.
[0007] [Figure 2] This is a flowchart of a multimodal process for predicting visual acuity responses, relating to various embodiments.
[0008] [Figure 3] This is a block diagram of a multimodal neural network system according to various embodiments.
[0009] [Figure 4]This is a flowchart of a first single-mode process for predicting visual acuity response according to various embodiments.
[0010] [Figure 5] This is a block diagram of a first single-mode neural network system according to various embodiments.
[0011] [Figure 6] This is a flowchart of a second single-mode process for predicting visual acuity response according to various embodiments.
[0012] [Figure 7] This is a block diagram of a second single-mode neural network system according to various embodiments.
[0013] [Figure 8] This is a block diagram of a computer system according to various embodiments.
[0014] It should be understood that the drawings are not necessarily drawn to a consistent scale, and the objects within the drawings are not necessarily drawn to a consistent scale with respect to each other. The drawings are intended to provide clarity and understanding of the various embodiments of the apparatus, systems, and methods disclosed herein. Wherever possible, the same reference numerals are used throughout the drawings to refer to the same or similar parts. Furthermore, it should be understood that the drawings are not in any way limited to the scope of this instruction. [Modes for carrying out the invention]
[0015] Detailed explanation overview Determining a subject's response to age-related macular degeneration (AMD) treatment may include determining the subject's visual acuity response (VAR). A subject's visual acuity is the sharpness of their vision, which can be measured by their ability to distinguish letters or numbers at a given distance. Visual acuity is often confirmed by visual acuity testing and measured according to the standard Snellen's visual acuity chart. However, other visual acuity measurements may be used instead of the Snellen's chart. Retinal images may provide information that can be used to estimate a subject's visual acuity. For example, color fundus (CF) images may be used to estimate the subject's visual acuity at the time the color fundus image was taken.
[0016] However, in certain cases, such as clinical trials, it may be desirable to be able to predict a subject's future visual acuity in response to an AMD treatment. For example, it may be desirable to predict whether a subject's visual acuity has improved over a selected period after treatment (e.g., 3, 6, 9, or 12 months after treatment). Furthermore, it may be desirable to classify any such predicted visual acuity improvement. Such predictions and classifications can enable treatment regimens that are individualized for a given subject. For example, predictions about a subject's visual response to a particular AMD treatment can be used to customize the treatment dose (e.g., injection dose), the interval at which the treatment (e.g., injection) is administered, or both. Moreover, such predictions can improve clinical trial screening, pre-screening, or both by enabling the exclusion of subjects predicted to be less responsive to the treatment.
[0017] Accordingly, the various embodiments described herein provide methods and systems for predicting visual acuity responses to AMD treatments. In particular, imaging data from one or more imaging modalities are received and processed by a neural network system to predict a visual acuity response (VAR) output. The VAR output may include a predicted change in visual acuity of the subject receiving treatment. In some cases, the VAR output corresponds to a predicted change in visual acuity, in that the VAR output may be further processed to determine this predicted change. Thus, the VAR output can be an indicator of the predicted change in visual acuity. In one or more embodiments, these different imaging modalities include color fundus imaging and / or optical coherence tomography (OCT).
[0018] Color fundus imaging is a two-dimensional imaging modality. It captures a field of view of approximately 30 to 50 degrees of the retina and optic nerve. In addition to being widely available and easy to use, color fundus imaging can capture the appearance of the optic nerve and the presence of intraocular blood accumulation better than other imaging modalities. However, color fundus imaging may not be able to capture thickness or volume data regarding the retina.
[0019] OCT can be considered a three-dimensional imaging modality. In particular, OCT can be used to capture images with a resolution on the order of micrometers (e.g., a resolution of up to about 10 μm, 9 μm, 8 μm, 7 μm, 6 μm, 5 μm, 4 μm, 3 μm, 2 μm, 1 μm, or more; at least about 1 μm, 2 μm, 3 μm, 4 μm, 5 μm, 6 μm, 7 μm, 8 μm, 9 μm, 10 μm, or less; or a resolution within a range defined by any two of the foregoing values) to provide depth information. OCT images can provide retinal thickness and / or volume information that cannot be confirmed, or cannot be easily or accurately confirmed, using color fundus imaging. For example, OCT images can be used to measure retinal thickness. Additionally, OCT images can be used to reveal and distinguish between fluid within the retina and subretinal fluid (e.g., subretinal fluid). Further, OCT images can be used to identify the location of abnormal new blood vessels within the eye. However, OCT images may not be as accurate in identifying blood accumulations as compared to color fundus images.
[0020] It is recognized that various embodiments provided herein can achieve sufficient accuracy, precision, and / or recall metrics for a neural network trained using only color fundus images or only OCT images to provide a highly reliable VAR prediction of the response to AMD treatment. Such a neural network can be particularly beneficial when only one of a color fundus image and an OCT image is available for a particular subject.
[0021] It is recognized that various embodiments provided herein are such that each of color fundus imaging and OCT can provide more accurate information regarding at least one retinal feature as compared to the other of these two imaging modalities. Thus, it is recognized that various embodiments described herein use the information provided by both of these different imaging modalities, which may enable an improved VAR prediction of the response to AMD treatment as compared to using each imaging modality independently. Such multimodal approaches may generally enable a faster, more efficient, and more accurate prediction of visual acuity response as compared to at least some of the currently available methodologies for predicting AMD treatment outcomes.
[0022] Recognizing and considering the importance and usefulness of methodologies and systems that can provide the improvements described above, this specification describes various embodiments for predicting VAR for AMD treatment. More specifically, this specification describes various embodiments of methods and systems for processing imaging data obtained via one or two different imaging modalities using a neural network system (e.g., a convolutional neural network system) to generate a VAR output that enables prediction of a subject's future visual acuity during a selected period after treatment.
[0023] Furthermore, this embodiment facilitates the creation of an individualized treatment regimen for an individual subject and ensures an appropriate dosage and / or interval between injections. In particular, the single-mode and multimodal approaches for predicting VAR presented herein can generate accurate, efficient, and / or appropriate individualized treatments and / or dosing schedules and can help to enhance clinical cohort selection and / or clinical trial design.
[0024] Definitions This disclosure is not limited to these exemplary embodiments and uses, or the ways in which these exemplary embodiments and uses operate or are described herein. Furthermore, figures may be simplified or partial, and the dimensions of elements in the figures may be exaggerated or disproportionate.
[0025] Furthermore, wherever the terms “on,” “attached to,” “connected to,” “coupled to,” or similar terms are used herein, one element (e.g., a component, material, layer, substrate, etc.) can be “on,” “attached to,” “connected to,” or “coupled to” another element, regardless of whether one element is directly on top of another element, directly attached to another element, connected to another element, or coupled to another element, or whether one or more intervening elements exist between one element and the other. Furthermore, wherever a list of elements (e.g., elements a, b, c) is referenced, such reference is intended to include any one of the enumerated elements, any combination of fewer elements than all of the enumerated elements, and / or all combinations of the enumerated elements. The division of sections herein is merely for the convenience of consideration and does not limit any combination of elements described.
[0026] The term “subject” may refer to a subject in a clinical trial, a person undergoing treatment, a person undergoing anti-cancer therapy, a person being monitored for remission or recovery, a person undergoing a preventive health analysis (e.g., due to their medical history), or any other person or patient of interest. In various cases, “subject” and “patient” may be used interchangeably herein.
[0027] Unless otherwise defined, scientific and technical terms used in connection with these instructions herein shall have meanings generally understood by those skilled in the art. Furthermore, unless otherwise required by context, singular terms shall include plural forms, and plural terms shall include singular forms. In general, nomenclature and techniques used in connection with chemistry, biochemistry, molecular biology, pharmacology, and toxicology are described herein, are well known and commonly used in the art.
[0028] As used herein, “substantially” means sufficient to function for the intended purpose. Therefore, the term “substantially” allows for minor, slight variations from absolute or perfect conditions, dimensions, measurements, results, etc., which are expected by those skilled in the art but do not significantly affect the overall performance. When used in relation to numerical values, or parameters or characteristics that can be expressed numerically, “substantially” means within 10 percent.
[0029] The term "plural" means two or more.
[0030] As used herein, the term “plural” may mean two, three, four, five, six, seven, eight, nine, ten or more.
[0031] As used herein, the term "set" means one or more items. For example, a set of items includes one or more items.
[0032] As used herein, the phrase “at least one of” means, when used with a list of items, that one or more different combinations of the enumerated items may be used, and only one of the items in the list may be required. An item can be a specific object, thing, step, action, process, or category. In other words, “at least one of” means that any combination or any number of items from the list may be used, but not all of the items in the list are required. For example, but not limited to, “at least one of item A, item B, or item C” means item A, item A and item B, item B, item A, item B, and item C, item B and item C, or item A and C. In some cases, “at least one of item A, item B, or item C” means, but not limited to, two of item A, one of item B and ten of item C, four of item B and seven of item C, or several other suitable combinations.
[0033] As used herein, the term "or" can have both disjunctive and conjunctive meanings. That is, the phrase "A or B" may refer to A alone, B alone, or both A and B.
[0034] In drawings, the same number refers to the same element.
[0035] As used herein, “Model” may include one or more algorithms, one or more mathematical techniques, one or more machine learning algorithms, or a combination thereof.
[0036] As used herein, “machine learning” includes the practice of using algorithms to analyze data, learn from it, and then make decisions or predictions about something in the world. Machine learning uses algorithms that can learn from data without relying on rule-based programming.
[0037] As used herein, “artificial neural network” or “neural network” (NN) may refer to a mathematical algorithm or computational model that mimics an interconnected group of artificial neurons that process information based on a connectivity-theoretic approach to computation. A neural network, sometimes also called a neural net, can use one or more layers of linear units, nonlinear units, or both, to predict the output of an incoming input according to a mathematical operation defined by parameters or weight coefficients determined in the training modes described herein. Some neural networks include one or more inner or hidden layers in addition to an output layer. The output of each inner or hidden layer may be used as an input to the next layer in the network, i.e., the next inner or hidden layer or output layer. Each layer of the network produces an output from an incoming input according to the current values of each set of parameters. In various embodiments, a reference to “neural network” may refer to one or more neural networks.
[0038] A neural network can process information in two ways: it is in training mode when it is being trained, and it is in inference (or prediction) mode when it actually performs what it has learned. A neural network can learn through a feedback process (e.g., backpropagation) that allows the network to adjust the weight coefficients of individual nodes in intermediate inner or hidden layers (correcting its behavior) so that the output matches the output of the training data. In other words, a neural network learns by being provided with training data (learning examples) and eventually learns how to arrive at the correct output even when presented with a new range or set of inputs. The set of mathematical operations, parameters, and / or weight coefficients learned during training mode may be referred to herein as the “trained model”. The trained model can then be applied to a new range or set of inputs in prediction mode. A neural network may include, but is not limited to, a feedforward neural network (FNN), a recurrent neural network (RNN), a modular neural network (MNN), a convolutional neural network (CNN), a fully convolutional neural network (FCN), a residual neural network (ResNet), an ordinary differential equation neural network (neural-ODE), a deep neural network, or at least one of the other types of neural networks.
[0039] Prediction of visual acuity response Figure 1 is a block diagram of the prediction system 100 according to various embodiments. The prediction system 100 is used to predict the visual acuity response (VAR) of one or more subjects in response to an AMD treatment. The AMD treatment may be an anti-VEGF treatment such as ranibizumab, which may be administered via intravitreal injection or another mode of administration, for example, but is not limited thereto.
[0040] The prediction system 100 includes a computing platform 102, data storage 104, and a display system 106. The computing platform 102 can take various forms. In one or more embodiments, the computing platform 102 includes a single computer (or computer system) or multiple computers communicating with each other. In other examples, the computing platform 102 takes the form of a cloud computing platform. In some examples, the computing platform 102 takes the form of a mobile computing platform (e.g., a smartphone, tablet, smartwatch, etc.).
[0041] The data storage 104 and the display system 106 each communicate with the computing platform 102. In some examples, the data storage 104, the display system 106, or both may be considered part of the computing platform 102, or otherwise integrated. Thus, in some examples, the computing platform 102, the data storage 104, and the display system 106 may be separate components that communicate with each other, while in other examples, some combination of these components may be integrated together.
[0042] The prediction system 100 includes a data analyzer 108, which may be implemented using hardware, software, firmware, or a combination thereof. In one or more embodiments, the data analyzer 108 is implemented on a computing platform 102. The data analyzer 108 uses a neural network system 112 to process one or more inputs 110 to predict (or generate) a visual acuity response (VAR) output 114. The VAR output 114 includes a predicted change in the visual acuity of a subject being treated. In some embodiments, one or more inputs 110 include a first input 110a and a second input 110b, as shown in Figure 1. Such embodiments may be referred to herein as “multimodal”. In some embodiments, one or more inputs 110 include a single input. Such embodiments may be referred to herein as “single-mode”.
[0043] The neural network system 112 may include any number or combination of neural networks. In one or more embodiments, the neural network system 112 takes the form of a convolutional neural network (CNN) system including one or more neural network subsystems. In some embodiments, at least one of these one or more neural network subsystems may be a convolutional neural network itself. In other embodiments, at least one of these one or more neural network subsystems may be a deep learning neural network (or deep neural network). In some embodiments, the neural network system 112 includes a multimodal neural network system as described herein with respect to Figure 3. In some embodiments, the neural network system 112 includes a first single-mode neural network system as described herein with respect to Figure 5. In some embodiments, the neural network system 112 includes a second single-mode neural network system as described herein with respect to Figure 7.
[0044] In a multimodal approach, the neural network system 112 may be trained through a single process in which various parts of the neural network system 112 are trained together (e.g., simultaneously). Therefore, in a multimodal approach, the neural network system 112 does not need to generate an output after the first training, integrate the output into the neural network system 112, and then perform a second training. In a multimodal approach, the entire neural network system 112 may be trained together (e.g., simultaneously), which can improve training efficiency and / or reduce the processing power required for this training.
[0045] Multimodal neural network Figure 2 is a flowchart of a multimodal process 200 for predicting visual acuity responses according to various embodiments. In one or more embodiments, the process 200 is implemented using the prediction system 100 described herein with respect to Figure 1.
[0046] Step 202 includes receiving a first input which includes two-dimensional imaging data associated with a subject undergoing a procedure (such as the AMD procedure described herein). The two-dimensional imaging data may take the form of color fundus imaging data associated with the subject undergoing the procedure. For example, the color fundus imaging data may be a color fundus image associated with the subject undergoing the procedure, or data extracted from such a color fundus image. The color fundus imaging data may be a color fundus image of the eye of the subject undergoing the procedure, or data extracted from that color fundus image.
[0047] Step 204 includes receiving a second input to the neural network system, which includes three-dimensional imaging data associated with the subject receiving treatment. The three-dimensional imaging data may include OCT imaging data, which may include data extracted from OCT images associated with the subject receiving treatment (e.g., OCT in-plane images), which may include tabular data extracted from such OCT images, or which may include other forms of such OCT imaging data. The OCT imaging data may, for example, take the form of OCT images associated with the subject receiving treatment. The OCT imaging data may be an OCT image of the eye of the subject receiving treatment or data extracted from such OCT image. In one or more embodiments, the second input includes other data associated with the subject receiving treatment, such as, for example, visual acuity measurement data associated with the subject receiving treatment, demographic statistics associated with the subject receiving treatment, or both. The visual acuity measurement data may include one or more visual acuity measurements associated with the subject receiving treatment (e.g., best corrected visual acuity (BCVA) measurements). The demographic statistics may include, for example, the age, sex, height, weight, or overall health level of the subject receiving treatment. In various embodiments, both visual acuity measurement data and demographic statistical data are baseline data associated with the subject receiving treatment.
[0048] In one or more embodiments, the second input takes the form of tabular data including BCVA measurements, demographic statistics, and three-dimensional imaging data (e.g., OCT thickness, OCT volume, etc.). Because OCT images are large and complex, converting these OCT images into tabular format can help a neural network system process the data contained in these images. In particular, by converting OCT imaging data into tabular format, the processing power and size of the part of the neural network system processing this tabular data can be reduced compared to processing the OCT images (e.g., in-plane OCT images). These processing savings may allow the second input to be more easily integrated with the first input.
[0049] Step 206 involves predicting a visual acuity response (VAR) output using a first and second input via a neural network system, the VAR output containing a predicted change in the visual acuity response of the treated subject. In some embodiments, the VAR output identifies the predicted change. In other embodiments, the VAR output corresponds to the predicted change, in that the VAR output may be further processed to determine the predicted change. The predicted VAR output may correspond to the initiation of AMD treatment or a selected period after administration. For example, the VAR output may allow prediction of the subject's visual acuity response for at least about 3 months, 6 months, 9 months, 12 months, 18 months, or 24 months or more after the initiation of treatment, up to about 24 months, 18 months, 12 months, 9 months, 6 months, or 3 months or less after the initiation of treatment, or for a period after the initiation of treatment within the range defined by any two of the aforementioned values.
[0050] In one or more embodiments, predicting a VAR output includes generating a first output using two-dimensional imaging data via a neural network system and generating a second output using three-dimensional imaging data via a neural network system. In some embodiments, the VAR output is generated by fusing the first and second outputs. That is, in some embodiments, the first output is generated using a first part of the neural network system (e.g., a first neural network subsystem described herein with respect to Figure 3), and the second output is generated using a second part of the neural network system (e.g., a second neural network subsystem described herein with respect to Figure 3). The first and second outputs can then be fused to form a fused input to a third part of the neural network system (e.g., a third neural network subsystem described herein with respect to Figure 3). The fused input can then be used by the third neural network subsystem to generate a VAR output that provides an index of the predicted change in the subject's visual acuity.
[0051] In some embodiments, the first output includes one or more features extracted from two-dimensional imaging data. In some embodiments, the second output includes one or more features extracted from three-dimensional imaging data. The features extracted from the two-dimensional imaging data and the features extracted from the three-dimensional imaging data can then be fused together to form a fused input. A third part of the neural network system can then generate a VAR output based on the fused input. In some embodiments, features extracted from 2D imaging data and / or 3D imaging data are associated with regions containing abnormalities (such as lesions, abnormal bleeding, scar tissue, and / or tissue atrophy) on or inside the subject's eye, the size of such regions, the perimeter of such regions, the area of such regions, the shape descriptive features of such regions, the distance of such regions to various features of the eye (e.g., fovea, macula, retina, sclera, or choroid), the continuity of such regions, wedge-shaped subretinal reduction, attenuation and destruction of retinal pigment epithelium (RPE), high-reflectivity focal points, reticular pseudodrusen (RPD), multilayer thickness reduction, photoreceptor atrophy, low-reflectivity core within drusen, high central drusen volume, previous visual acuity, extraretinal fallopian tube formation, choroidal capillary lamellar fluid space, coloration of any region of the 2D imaging data and / or 3D imaging data, fading of any region of the 2D imaging data and / or 3D imaging data, or any combination thereof.
[0052] In some embodiments, the first and second outputs are fused to form an integrated multi-channel input that can then undergo a subsequent feature extraction process by a third part of the neural network system. The features extracted by the feature extraction process can then be used as a basis for generating a VAR output. Features extracted by the feature extraction process (and / or fused input) may include, or be associated with, areas containing abnormalities (such as lesions, abnormal bleeding, scar tissue, and / or tissue atrophy) on or inside the subject's eye, the size of such areas, the perimeter of such areas, the area of such areas, descriptive features of the shape of such areas, the distance of such areas to various features of the eye (e.g., fovea, macula, retina, sclera, or choroid), the continuity of such areas, wedge-shaped subretinal reduction, attenuation and destruction of retinal pigment epithelium (RPE), high-reflectivity focus, reticular pseudodrusen (RPD), multilayer thickness reduction, photoreceptor atrophy, low-reflectivity core within drusen, high central drusen volume, previous visual acuity, outer retinal canal formation, choroidal capillary lamellar fluid space, coloration of 2D imaging data and / or 3D imaging data or any region thereof, fading of 2D imaging data and / or 3D imaging data or any region thereof, or any prior combination thereof.
[0053] In various embodiments, the VAR output is a value or score that identifies a predicted change in a subject's visual acuity. For example, the VAR output may be a value or score that classifies the subject's visual acuity response in terms of the predicted level of improvement (e.g., improvement character) or decline (e.g., vision loss). In one specific example, the VAR output may be a predicted numerical change in BCVA that is later processed and identified as belonging to one of several different classes of BCVA change, each class of BCVA corresponding to a different range of improvement characters. In another example, the VAR output may be the predicted change class itself. In yet another example, the VAR output may be a predicted change in some other measure of visual acuity.
[0054] In other embodiments, the VAR output may be a value or expression output that requires one or more additional processing steps to reach the predicted change in visual acuity. For example, the VAR output may be the predicted future BCVA of the subject during a post-treatment period (e.g., at least about 3 months, 6 months, 9 months, 12 months, 18 months, 24 months or more after treatment, up to about 24 months, 18 months, 12 months, 9 months, 6 months, 3 months or less after treatment, or within the range defined by any two of the aforementioned values). One or more additional processing steps may include calculating the difference between the predicted future BCVA and the baseline BCVA to determine the predicted change in visual acuity.
[0055] In some embodiments, the method further includes training a neural network system before receiving first and second inputs. In some embodiments, the neural network system is trained using two-dimensional data associated with a first plurality of subjects that have been previously treated and three-dimensional data associated with a second plurality of subjects that have been previously treated. The first and second plurality can be any number of subjects, e.g., at least about 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, 900,000, 1,000,000 or more subjects, up to about 1,000,000, 900,000, This may include data associated with the number of subjects up to or below 800,000, 700,000, 600,000, 500,000, 400,000, 300,000, 200,000, 100,000, 90,000, 80,000, 70,000, 60,000, 50,000, 40,000, 30,000, 20,000, 10,000, 9,000, 8,000, 7,000, 6,000, 5,000, 4,000, 3,000, 2,000, or 1,000, or the number of subjects within the range defined by any two of the aforementioned values.
[0056] In some embodiments, the first and second multiples are the same; that is, in some cases, the first and second multiples include exactly the same subjects. In some embodiments, the first and second multiples are different; that is, in some cases, the first multiple includes one or more subjects not characterized in the second multiple, and vice versa. In some embodiments, the first and second multiples partially overlap; that is, in some cases, one or more subjects are characteristic in both the first and second multiples.
[0057] In some embodiments, training the neural network system further includes using visual acuity measurements associated with a second group of subjects who have previously been treated, demographic statistics associated with the second group, or a combination thereof.
[0058] In some embodiments, the neural network system is trained using focus loss, cross-entropy loss, or weighted cross-entropy loss.
[0059] Figure 3 is a block diagram of the multimodal neural network system 300. In some embodiments, the multimodal neural network system is configured for use with the prediction system 100 described herein with respect to Figure 1. In some embodiments, the multimodal neural network system is configured to implement the method 200 (or any of steps 202, 204, and 206) described herein with respect to Figure 2.
[0060] In some embodiments, the multimodal neural network system comprises a first neural network subsystem 310. In some embodiments, the first neural network subsystem includes at least one first input layer 312 and at least one first high-density inner layer 314. In some embodiments, the first input layer is configured to receive a first input as described herein with respect to Figure 2. In some embodiments, at least one first high-density inner layer is configured to apply a first trained model to the first input layer.
[0061] In the illustrated example, at least one first high-density inner layer includes a trained image recognition model 314a and at least one output high-density inner layer 314b. In some embodiments, the trained image recognition model is configured to apply the image recognition model to the first input layer. In some embodiments, the image recognition model includes a pre-trained image recognition model. In some embodiments, the pre-trained image recognition model includes a deep residual network such as ResNet-34, ResNet-50, ResNet-101, or ResNet-152.
[0062] In some embodiments, the output high-density inner layer receives the output from the image recognition model and applies additional operations to the output from the image recognition model. In some embodiments, the additional operations are learned during the training of the first trained model. In some embodiments, the image recognition model is not updated during the training of the first trained model. In some embodiments, the output high-density inner layer is configured to apply mean pooling and / or softmax activation.
[0063] Although Figure 3 shows a single output high-density inner layer, at least one output high-density inner layer may include any number of high-density inner layers. In some embodiments, at least one output high-density inner layer includes at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 or more high-density inner layers, up to about 100, 90, 80, 70, 60, 50, 40, 30, 20, 10, 9, 8, 7, 6, 5, 4, 3, 2, or one high-density inner layer, or a number of high-density inner layers within the range defined by any two of the aforementioned values. Each output high-density inner layer may be configured to apply mean pooling, normalized linear (ReLU) activation, and / or softmax activation.
[0064] In some embodiments, the multimodal neural network system comprises a second neural network subsystem 320. In some embodiments, the second neural network subsystem includes at least one second input layer 322 and at least one second high-density inner layer 324. In some embodiments, the second input layer is configured to receive a second input as described herein with respect to Figure 2. In some embodiments, at least one second high-density inner layer is configured to apply a second trained model to the second input layer.
[0065] In the illustrated example, at least one second high-density inner layer includes three high-density inner layers 324a, 324b, and 324c. In some embodiments, high-density inner layer 324a is configured to apply a first set of operations to a second input layer. In some embodiments, high-density inner layer 324b is configured to apply a second set of operations to high-density inner layer 324a. In some embodiments, high-density inner layer 324c is configured to apply a third set of operations to high-density inner layer 324b. In some embodiments, the first, second, and third sets of operations are learned during training of a second trained model. In some embodiments, high-density inner layers 324a and 324b are configured to apply ReLU activation, and high-density inner layer 324c is configured to apply softmax activation.
[0066] Although Figure 3 shows a configuration including three second high-density inner layers, at least one second high-density inner layer may include any number of high-density inner layers. In some embodiments, at least one second high-density inner layer includes at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 or more high-density inner layers, at most about 100, 90, 80, 70, 60, 50, 40, 30, 20, 10, 9, 8, 7, 6, 5, 4, 3, 2, or one high-density inner layer, or several high-density inner layers within the range defined by any two of the aforementioned values. Each of the second high-density inner layers may be configured to apply mean pooling, normalized linear (ReLU) activation, and / or softmax activation.
[0067] In some embodiments, the multimodal neural network system comprises a third neural network subsystem 330. In some embodiments, the third neural network subsystem includes at least one third high-density inner layer 332. In some embodiments, the third at least one third high-density inner layer is configured to receive a first output from at least a first high-density inner layer associated with the first neural network subsystem and a second output from at least a second high-density inner layer associated with the second neural network subsystem.
[0068] In the illustrated example, at least one third high-density inner layer includes a single layer. In some embodiments, the single layer is configured to apply a set of operations to the first and second outputs. In some embodiments, the set of operations is learned during the training of the third trained model. In some embodiments, the third high-density inner layer is configured to apply softmax activation.
[0069] Although Figure 3 shows a single third high-density inner layer, at least one third high-density inner layer may include any number of high-density inner layers. In some embodiments, at least one third high-density inner layer includes at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 or more high-density inner layers, at most about 100, 90, 80, 70, 60, 50, 40, 30, 20, 10, 9, 8, 7, 6, 5, 4, 3, 2, or one high-density inner layer, or several high-density inner layers within the range defined by any two of the aforementioned values. Each of the third high-density inner layers may be configured to apply mean pooling, normalized linear (ReLU) activation, and / or softmax activation.
[0070] In some embodiments, the neural network system is configured to output classification data 340. In some embodiments, the classification data includes a first likelihood 342 that a treated subject may achieve a score of less than 5 characters in a visual acuity test during the post-treatment period, a second likelihood 344 that a treated subject may achieve a score of 5 to 9 characters, a third likelihood 346 that a treated subject may achieve a score of 10 to 14 characters, and / or a fourth likelihood 348 that a treated subject may achieve a score of more than 15 characters. In some embodiments, the output classification data is placed as the output layer of the neural network system.
[0071] Although Figure 3 shows the data as including four classes, the classification data may include any number of classes. For example, the classification data may include at least approximately 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more classes, and at most approximately 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, or two classes, or several classes within the range defined by any two of the aforementioned values. For example, the classification data may include first and second likelihoods that a treated subject is likely to achieve a score of less than 10 characters and a score of more than 11 characters, respectively. As a further example, the classification data may include the first, second, third, fourth, fifth, sixth, seventh, eighth, ninth, tenth, and eleventh likelihoods that a treated subject is likely to achieve a score of less than two characters, two to three characters, four to five characters, six to seven characters, eight to nine characters, ten to eleven characters, twelve to thirteen characters, fourteen to fifteen characters, sixteen to seventeen characters, eighteen to nineteen characters, and more than twenty characters, respectively. Those skilled in the art will recognize that many variations are possible.
[0072] In some embodiments, the first, second, and third trained models are trained together. In some embodiments, the first, second, and third trained models are trained simultaneously. For example, in some embodiments, training data in the form of two-dimensional imaging data associated with a first group of previously treated subjects is provided to a first neural network subsystem, and at the same time, training data in the form of three-dimensional imaging data associated with a first group of previously treated subjects is provided to a second neural network subsystem. Then, the first, second, and third models associated with the first, second, and third neural network subsystems, respectively, are trained simultaneously. In this way, a multimodal neural network system can be trained end-to-end without requiring separate, independent, or sequential training of its components.
[0073] In some embodiments, the neural network system is configured to apply an exemplary attention gate mechanism.
[0074] Single-mode neural network using 2D data Figure 4 is a flowchart of a first single-mode process 400 for predicting visual acuity response according to various embodiments. In one or more embodiments, the process 400 is implemented using the prediction system 100 described herein with respect to Figure 1.
[0075] Step 402 includes receiving an input containing two-dimensional imaging data associated with a subject undergoing a procedure (such as the AMD procedure described herein). The two-dimensional imaging data may take the form of any two-dimensional imaging data described herein (e.g., any two-dimensional imaging data described herein with respect to Figure 1, Figure 2, or Figure 3).
[0076] Step 404 involves predicting a visual acuity response (VAR) output using the input via a neural network system, the VAR output comprising the predicted change in the visual acuity response of the subject being treated. In some embodiments, the VAR output includes any VAR output described herein (e.g., any VAR output described herein with respect to Figure 1, Figure 2, or Figure 3).
[0077] In some embodiments, the method further includes training the neural network system before receiving first and second inputs. In some embodiments, the neural network system is trained using two-dimensional data associated with a plurality of subjects that have previously been treated. The plurality may be at least about 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 20,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, 900,000, 1,000,000, or more subjects, up to about 1,000,000, 900,000, 8600,000, 500,000 This may include data associated with any number of subjects, such as the number of subjects in the range defined by any two of the aforementioned values, such as 400,000, 300,000, 200,000, 100,000, 90,000, 80,000, 70,000, 60,000, 50,000, 40,000, 30,000, 20,000, 10,000, 9,000, 8,000, 7,000, 6,000, 5,000, 4,000, 3,000, 2,000, 1,000, or less, or the number of subjects within the range defined by any two of the aforementioned values.
[0078] Figure 5 is a block diagram of the first single-mode neural network system 500. In some embodiments, the first single-mode neural network system is configured for use with the prediction system 100 described herein with respect to Figure 1. In some embodiments, the first single-mode neural network system is configured to implement the method 400 (or either steps 402 and 404) described herein with respect to Figure 4.
[0079] In some embodiments, the first single-mode neural network system includes at least one input layer 502 and at least one high-density inner layer 504. In some embodiments, the input layer is configured to receive the input described herein with respect to Figure 4. In some embodiments, at least one high-density inner layer is configured to apply a trained model to the input layer.
[0080] In the illustrated example, at least one high-density inner layer includes a trained image recognition model 504a and at least one output high-density inner layer 504b. In some embodiments, the trained image recognition model is configured to apply the image recognition model to the input layer. In some embodiments, the image recognition model includes any image recognition model described herein (for example, any image recognition model described herein with respect to Figure 3).
[0081] In some embodiments, the output high-density inner layer receives the output from the image recognition model and applies additional operations to the output from the image recognition model. In some embodiments, the additional operations are learned during the training of the trained model. In some embodiments, the image recognition model is not updated during the training of the trained model. In some embodiments, the output high-density inner layer is configured to apply mean pooling and / or softmax activation.
[0082] Although Figure 5 shows a single output high-density inner layer, at least one output high-density inner layer may include any number of high-density inner layers. In some embodiments, at least one output high-density inner layer includes at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 or more high-density inner layers, at most about 100, 90, 80, 70, 60, 50, 40, 30, 20, 10, 9, 8, 7, 6, 5, 4, 3, 2, or one high-density inner layer, or a number of high-density inner layers within the range defined by any two of the aforementioned values. Each output high-density inner layer may be configured to apply mean pooling, normalized linear (ReLU) activation, and / or softmax activation.
[0083] In some embodiments, the neural network system is configured to output classification data 510. In some embodiments, the classification data includes a first likelihood 512 that a treated subject may achieve a score of less than 5 characters in a visual acuity test during the post-treatment period, a second likelihood 514 that a treated subject may achieve a score of 5 to 9 characters, a third likelihood 516 that a treated subject may achieve a score of 10 to 14 characters, and / or a fourth likelihood 518 that a treated subject may achieve a score of more than 15 characters. In some embodiments, the output classification data is placed as the output layer of the neural network system.
[0084] Although Figure 5 shows data containing four classes, the classification data may contain any number of classes, as described herein (for example, as described herein with respect to Figure 3).
[0085] In some embodiments, the neural network system is configured to apply an exemplary attention gate mechanism.
[0086] Single-mode neural network using 3D data Figure 6 is a flowchart of a second single-mode process 600 for predicting visual acuity response according to various embodiments. In one or more embodiments, the process 600 is implemented using the prediction system 100 described herein with respect to Figure 1.
[0087] Step 602 includes receiving an input to a neural network system that includes three-dimensional imaging data associated with the subject being treated. The three-dimensional imaging data may include any three-dimensional imaging data described herein (for example, any three-dimensional imaging data described herein with respect to Figure 1, Figure 2, or Figure 3).
[0088] Step 604 involves predicting a visual acuity response (VAR) output using the input via a neural network system, the VAR output comprising the predicted change in the visual acuity response of the subject being treated. In some embodiments, the VAR output includes any VAR output described herein (e.g., any VAR output described herein with respect to Figure 1, Figure 2, or Figure 3).
[0089] In some embodiments, the method further includes training a neural network system before receiving first and second inputs. In some embodiments, the neural network system is trained using three-dimensional data associated with a plurality of subjects that have previously been treated. The plurality may be at least about 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 20,000, 30,000, 40,000, 50,000, 60,000, 70,000, 800,000, 900,000, 1,000,000, or more subjects, up to about 1,000,000, 900,000, 800,000, 70600,000, 500,000 This may include data associated with any number of subjects, such as the number of subjects in the range defined by any two of the aforementioned values, such as 400,000, 300,000, 200,000, 100,000, 90,000, 80,000, 70,000, 60,000, 50,000, 40,000, 30,000, 20,000, 10,000, 9,000, 8,000, 7,000, 6,000, 5,000, 4,000, 3,000, 2,000, 1,000, or less, or the number of subjects within the range defined by any two of the aforementioned values.
[0090] Figure 7 is a block diagram of the second single-mode neural network system 700. In some embodiments, the second single-mode neural network system is configured for use with the prediction system 100 described herein with respect to Figure 1. In some embodiments, the second single-mode neural network system is configured to implement the method 600 (or either step 602 and 604) described herein with respect to Figure 6.
[0091] In some embodiments, the second single-model neural network system includes at least one input layer 702 and at least one high-density inner layer 704. In some embodiments, the input layer is configured to receive the input described herein with respect to Figure 6. In some embodiments, at least one high-density inner layer is configured to apply a trained model to the input layer.
[0092] In the illustrated example, at least one high-density inner layer includes three high-density inner layers 704a, 704b, and 704c. In some embodiments, high-density inner layer 704a is configured to apply a first set of operations to the input layer. In some embodiments, high-density inner layer 704b is configured to apply a second set of operations to high-density inner layer 704a. In some embodiments, high-density inner layer 704c is configured to apply a third set of operations to high-density inner layer 704b. In some embodiments, the first, second, and third sets of operations are learned during training of the trained model. In some embodiments, high-density inner layers 704a and 704b are configured to apply ReLU activation, and high-density inner layer 704c is configured to apply softmax activation.
[0093] Although Figure 7 shows a configuration including three high-density inner layers, at least one high-density inner layer may include any number of high-density inner layers. In some embodiments, at least one high-density inner layer includes at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 or more high-density inner layers, up to about 100, 90, 80, 70, 60, 50, 40, 30, 20, 10, 9, 8, 7, 6, 5, 4, 3, 2, or one high-density inner layer, or several high-density inner layers within the range defined by any two of the aforementioned values. Each high-density inner layer may be configured to apply mean pooling, normalized linear (ReLU) activation, and / or softmax activation.
[0094] In some embodiments, the neural network system is configured to output classification data 710. In some embodiments, the classification data includes a first likelihood 712 that a treated subject may achieve a score of less than 5 characters in a visual acuity test during the post-treatment period, a second likelihood 714 that a treated subject may achieve a score of 5 to 9 characters, a third likelihood 716 that a treated subject may achieve a score of 10 to 14 characters, and / or a fourth likelihood 718 that a treated subject may achieve a score of more than 15 characters. In some embodiments, the output classification data is placed as the output layer of the neural network system.
[0095] Although Figure 7 shows data containing four classes, the classification data may contain any number of classes, as described herein (for example, as described herein with respect to Figure 3).
[0096] In some embodiments, the neural network system is configured to apply an exemplary attention gate mechanism.
[0097] In some embodiments, the systems and methods described herein are used to provide treatment recommendations. For example, in some embodiments, a neural network system is configured to generate a treatment output based on a VAR output. In some embodiments, the treatment output indicates a predicted change in the subject's visual acuity in response to the treatment. In some embodiments, the treatment recommendation is provided to a healthcare provider based on the treatment output. In some embodiments, the treatment recommendation prompts the healthcare provider to administer the treatment to the subject, depending on whether the treatment output is an improvement in the subject's visual acuity. In some embodiments, the step of administering the treatment includes intravitreal administration of the treatment or a derivative thereof at a treatment dose. In some embodiments, the treatment is ranibizumab, and the treatment dose is 0.3 milligrams (mg) or 0.5 mg. [Examples]
[0098] Examples Example 1: Prediction of visual acuity response in CATT test We developed a deep learning (DL) model to predict the visual acuity response (VAR) to ranibizumab (RBZ) using baseline (BL) characteristics and color fundus images (CFI) from patients with neovascular age-related macular degeneration. The VAR was formulated as a four-class classification problem (Class 1 = <5 letters, Class 2 = 5-9 letters, Class 3 = 10-14 letters, Class 4 = ≥15 letters). Each class was assigned based on the change in best corrected visual acuity (BCVA) from BL to 12 months. To solve the classification problem, we designed three DL models to process data from different modalities (two-dimensional and three-dimensional imaging modalities as described herein). Two different single-mode models were trained (as described herein with respect to Figures 4 and 5, and Figures 6 and 7, respectively) to process BL characteristics including BCVA, age, and CFI or optical coherence tomography (OCT) imaging biomarkers. The third model fused two subnetworks to generate the final classification, as described herein with respect to Figures 2 and 3. Exemplary attention mechanisms were used to enhance relevant parts of the input data and improve model performance. The data was split into training, validation, and test sets in a 3:1:1 ratio. Table 1 shows the loss type, number of epochs, and optimizer used during training for each model. [Table 1]
[0099] This study was a retrospective analysis of BL data from 284 patients who received monthly RBZ treatment in the randomized controlled trial of the Age-Related Macular Degeneration Treatment Trial (CATT) (NCT00593450). The CATT trial aimed to evaluate the relative efficacy and safety of RBZ and bevacizumab in monthly and as-needed regimens. The distribution across the four classes was unbalanced, with 64, 43, 52, and 125 patients in classes 1, 2, 3, and 4, respectively. Performance was assessed based on validation (N=56) and trial (N=57) data subsets using precision and area under the receiver operating characteristic (AUROC) curve. In addition, macro F1 (mF1) scores, class-specific F1 scores, and precision-recall (AUCPR) area under the curve were calculated to provide a more useful assessment of model performance.
[0100] Table 2 shows the various performance metrics for the three models. The performance metrics differed considerably among the three models (for example, the mF1 scores on the test dataset were 0.332, 0.236, and 0.354 for the OCT, CFI, and multimodal models, respectively). Furthermore, the results for each individual class showed significant variability, reflecting the presence of strong class imbalances in the data. [Table 2]
[0101] Table 3 shows the performance of the three models on a subset of trial data including the group that received monthly RBZ injections. Results are shown for the models with and without the exemplary attention mechanism. Table 4 shows the performance of the three models on a subset of trial data including all trial arms without the exemplary attention mechanism. [Table 3] [Table 4]
[0102] As shown in Tables 1-4, the multimodal model outperformed the CFI and, to a lesser extent, the OCT model across many performance metrics. However, for specific performance metrics, the CFI or OCT model provided the best performance. Therefore, all three models presented herein may be useful depending on the specific problem of interest.
[0103] Computerized Systems Figure 8 is a block diagram of a computer system according to various embodiments. The computer system 800 may be an example of an implementation of the computing platform 102 described in Figure 1. In one or more embodiments, the computer system 800 may include a bus 802 or other communication mechanism for communicating information and a processor 804 coupled to the bus 802 for processing information. In various embodiments, the computer system 800 may also include memory, which may be random access memory (RAM) 806 or other dynamic storage device, coupled to the bus 802 for determining instructions to be executed by the processor 804. The memory may also be used to store temporary variables or other intermediate information during the execution of instructions executed by the processor 804. In various embodiments, the computer system 800 may further include read-only memory (ROM) 808 or other static storage device coupled to the bus 802 for storing static information and instructions for the processor 804. A storage device 810, such as a magnetic disk or optical disk, may be provided and coupled to the bus 802 for storing information and instructions.
[0104] In various embodiments, the computer system 800 may be coupled via bus 802 to a display 812, such as a cathode ray tube (CRT) or liquid crystal display (LCD), to display information to the computer user. An input device 814, including alphanumeric and other keys, may be coupled to bus 802 to communicate information and command selections to the processor 804. Another type of user input device is a cursor control device 816, such as a mouse, joystick, trackball, gesture input device, gaze-based input device, or cursor direction keys, for communicating directional information and command selections to the processor 804 and controlling cursor movement on the display 812. This input device 814 typically has two degrees of freedom with two axes, a first axis (e.g., x) and a second axis (e.g., y), allowing the device to specify a position in a plane. However, it should be understood that input devices 814 enabling three-dimensional (e.g., x, y, and z) cursor movement are also contemplated herein.
[0105] In accordance with a particular implementation of this teaching, the results may be provided by the computer system 800 in response to a processor 804 executing one or more sequences of one or more instructions contained in RAM 806, or in response to a dedicated processing unit executing one or more sequences of one or more instructions contained in the dedicated RAM of those dedicated processing units. Such instructions may be read into RAM 806 from another computer-readable medium or computer-readable storage medium, such as a storage device 810. The execution of the instruction sequence contained in RAM 806 causes the processor 804 to execute the process described herein. Alternatively, hardwired circuits may be used instead of, or in combination with, software instructions to implement this teaching. Thus, implementations of this teaching are not limited to a particular combination of hardware circuits and software.
[0106] As used herein, the terms “computer-readable medium” (e.g., datastore, data storage, memory device, data storage device, etc.) or “computer-readable storage medium” refer to any medium involved in providing instructions to processor 804 for execution. Such mediums can take many forms, including but not limited to non-volatile media, volatile media, and transmission media. Examples of non-volatile media include, but are not limited to, optical, solid, and magnetic disks, such as memory device 810. Examples of volatile media include, but are not limited to, dynamic memory, such as RAM 806. Examples of transmission media include, but are not limited to, coaxial cables, copper wires, and optical fibers, including wires with bus 802.
[0107] Common forms of computer-readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tapes, or any other magnetic media, CD-ROMs, any other optical media, punch cards, paper tapes, any other physical media having a pattern of holes, RAM, PROMs, and EPROMs, flash EPROMs, any other memory chips or cartridges, or any other tangible media that a computer can read.
[0108] In addition to computer-readable media, instructions or data may be provided as signals on a transmission medium included in a communication device or system to provide a sequence of one or more instructions to a processor 804 of a computer system 800 for execution. For example, a communication device may include a transceiver having signals indicating instructions and data. The instructions and data are configured to cause one or more processors to implement the functions outlined in this disclosure. Typical examples of data communication transmission connections, but not limited to these, may include telephone modem connections, wide area networks (WANs), local area networks (LANs), infrared data connections, NFC connections, optical communication connections, and the like.
[0109] It should be understood that the methods, flowcharts, figures, and accompanying disclosures described herein may be implemented using the computer system 800 as a standalone device or on a distributed network of shared computing resources, such as a cloud computing network.
[0110] The methods described herein can be implemented by various means depending on the application. For example, these methodologies can be implemented in hardware, firmware, software, or any combination thereof. In the case of hardware implementation, the processing unit may be implemented in one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), processors, graphical processing units (GPUs), tensor processing units (TPUs), artificial intelligence (AI) accelerator ASICs, controllers, microcontrollers, microprocessors, electronic devices, other electronic units designed to perform the functions described herein, or a combination thereof.
[0111] In various embodiments, the methods described herein may be implemented as firmware and / or software programs and applications written in conventional programming languages such as C, C++, and Python. When implemented as firmware and / or software, the embodiments described herein may be implemented on a non-temporary computer-readable medium on which a program causing a computer to perform the methods described herein is stored. It should be understood that the various engines described herein may be provided on a computer system such as computer system 800, thereby the processor 804 performs the analysis and decision-making provided by these engines in accordance with instructions provided by one or a combination thereof of the memory components RAM 806, ROM 808, or storage devices 810, and user input provided via input device 814.
[0112] conclusion While this instruction is described in relation to various embodiments, it is not intended to be limited to such embodiments. On the contrary, this instruction includes various substitutions, modifications, and equivalents, as will be understood by those skilled in the art.
[0113] For example, the flowcharts and block diagrams described above illustrate the architecture, functionality, and / or operation of possible implementations of various methods and system embodiments. Each block in a flowchart or block diagram may represent a module, segment, function, operation or step, or a combination thereof. In some alternative implementations of an embodiment, one or more functions described in a block may be performed in an order different from the order shown in the diagram. For example, in some cases, two blocks shown consecutively may be executed substantially simultaneously or integrated in some way. In other cases, blocks may be executed in reverse order. Furthermore, in some cases, one or more blocks may be added to replace or supplement one or more other blocks in the flowchart or block diagram.
[0114] Accordingly, when describing various embodiments, this specification may present methods and / or processes as a specific set of steps. However, unless a method or process relies on a specific sequence of steps described herein, the method or process should not be limited to the specific sequence of steps described herein, and those skilled in the art will readily understand that the order may vary and still remain within the spirit and scope of various embodiments.
[0115] List of embodiments Embodiment 1. A method for predicting a visual acuity response, Receiving a first input which includes two-dimensional imaging data associated with the subject receiving treatment, Receiving a second input, which includes 3D imaging data associated with the subject undergoing treatment, A method comprising predicting a visual acuity response (VAR) output, including a predicted change in the visual acuity of a subject receiving a treatment in response to the treatment, using a neural network system with first and second inputs.
[0116] Embodiment 2. The method according to Embodiment 1, wherein the 3D imaging data includes optical coherence tomography (OCT) imaging data associated with the subject receiving treatment, and the 2D imaging data includes color fundus imaging data associated with the subject receiving treatment.
[0117] Embodiment 3. The method according to Embodiment 1 or 2, wherein the second input further includes visual acuity measurements associated with the subject receiving treatment and demographic data associated with the subject receiving treatment.
[0118] Embodiment 4. Predicting VAR output via a neural network system, To generate a first output using 2D imaging data associated with the subject receiving treatment, To generate a second output using 3D imaging data associated with the subject receiving the treatment, The method according to any one of embodiments 1 to 3, comprising generating a VAR output through the fusion of a first output and a second output.
[0119] Embodiment 5. A neural network system, A first neural network subsystem comprising at least one first input layer and at least one first high-density inner layer, wherein at least one first input layer is configured to receive a first input, and at least one first high-density inner layer is configured to apply a first trained model to the first input layer, A second neural network subsystem comprising at least one second input layer and at least one second high-density inner layer, wherein at least one second input layer is configured to receive a first input, and at least one second high-density inner layer is configured to apply a second trained model to the second input layer, The method according to any one embodiment of 1 to 4, further comprising: a third neural network subsystem including at least one third high-density inner layer configured to receive a first output from at least one first high-density inner layer and a second output from at least one second high-density layer, and to apply a third trained model to the first and second outputs to thereby predict a VAR output.
[0120] Embodiment 6. The method according to Embodiment 5, wherein at least one first high-density inner layer includes a trained image recognition model and an output high-density inner layer, and at least one second high-density inner layer includes a plurality of second high-density inner layers.
[0121] Embodiment 7. The method according to any one of Embodiments 1 to 6, further comprising training a neural network system using two-dimensional imaging data associated with a first group of previously treated subjects and three-dimensional imaging data associated with a second group of previously treated subjects, prior to receiving a first input and a second input.
[0122] Embodiment 8. The method according to Embodiment 7, further comprising using visual acuity measurements associated with a second group of subjects who have previously received treatment, demographic statistics associated with a second group of subjects who have previously received treatment, or a combination thereof, to train the neural network system.
[0123] Embodiment 9. A system for predicting visual acuity response, Non-temporary memory and One or more processors coupled to non-temporary memory and configured to read instructions from non-temporary memory to perform an operation on the system, wherein the operation is, Receiving a first input which includes two-dimensional imaging data associated with the subject receiving treatment, Receiving a second input, which includes 3D imaging data associated with the subject undergoing treatment, A system comprising one or more processors, which include predicting a visual acuity response (VAR) output, including a predicted change in the visual acuity of a subject receiving treatment in response to the treatment, using a first input and a second input via a neural network system.
[0124] Embodiment 10. The system according to Embodiment 9, wherein the 3D imaging data includes optical coherence tomography (OCT) imaging data associated with the subject receiving treatment, and the 2D imaging data includes color fundus imaging data associated with the subject receiving treatment.
[0125] Embodiment 11. The system according to Embodiment 9 or 10, wherein the second input further includes visual acuity measurements associated with the subject receiving treatment and demographic statistics associated with the subject receiving treatment.
[0126] Embodiment 12. Predicting VAR output via a neural network system, To generate a first output using 2D imaging data associated with the subject receiving treatment, To generate a second output using 3D imaging data associated with the subject receiving the treatment, A system according to any one of embodiments 9 to 11, comprising generating a VAR output through the fusion of a first output and a second output.
[0127] Embodiment 13. A neural network system, A first neural network subsystem comprising at least one first input layer and at least one first high-density inner layer, wherein at least one first input layer is configured to receive a first input, and at least one first high-density inner layer is configured to apply a first trained model to the first input layer, A second neural network subsystem comprising at least one second input layer and at least one second high-density inner layer, wherein at least one second input layer is configured to receive a first input, and at least one second high-density inner layer is configured to apply a second trained model to the second input layer, The system according to any one embodiment 9 to 12, comprising: a third neural network subsystem including at least one third high-density inner layer configured to receive a first output from at least one first high-density inner layer and a second output from at least one second high-density layer, and to apply a third trained model to the first and second outputs to predict a VAR output.
[0128] Embodiment 14. The system according to Embodiment 13, wherein at least one first high-density inner layer includes a trained image recognition model and an output high-density inner layer, and at least one second high-density inner layer includes a plurality of second high-density inner layers.
[0129] Embodiment 15. The system according to any one of Embodiments 9 to 14, further comprising training a neural network system using two-dimensional imaging data associated with a first group of previously treated subjects and three-dimensional imaging data associated with a second group of previously treated subjects, prior to receiving a first input and a second input.
[0130] Embodiment 16. The system according to Embodiment 15, further comprising using visual acuity measurements associated with a second group of subjects who have previously received treatment, demographic statistics associated with a second group of subjects who have previously received treatment, or a combination thereof, to train the neural network system.
[0131] Embodiment 17. A non-temporary machine-readable medium storing machine-readable instructions that can be executed to cause a system to perform an action, wherein the action is Receiving a first input which includes two-dimensional imaging data associated with the subject receiving treatment, Receiving a second input, which includes 3D imaging data associated with the subject undergoing treatment, A non-temporal machine-readable medium comprising: predicting a visual acuity response (VAR) output, including a predicted change in the visual acuity of a subject receiving treatment in response to the treatment, using a neural network system with first and second inputs.
[0132] Embodiment 18. A non-temporary machine-readable medium according to Embodiment 17, wherein the 3D imaging data includes optical coherence tomography (OCT) imaging data associated with the subject receiving treatment, and the 2D imaging data includes color fundus imaging data associated with the subject receiving treatment.
[0133] Embodiment 19. A non-temporary machine-readable medium according to Embodiment 17 or 18, wherein the second input further includes visual acuity measurements associated with the subject receiving treatment and demographic statistics associated with the subject receiving treatment.
[0134] Embodiment 20. Predicting VAR output via a neural network system, To generate a first output using 2D imaging data associated with the subject receiving treatment, To generate a second output using 3D imaging data associated with the subject receiving the treatment, A non-temporary machine-readable medium according to any one of embodiments 17 to 19, comprising generating a VAR output through the fusion of a first output and a second output.
[0135] Embodiment 21. A neural network system, A first neural network subsystem comprising at least one first input layer and at least one first high-density inner layer, wherein at least one first input layer is configured to receive a first input, and at least one first high-density inner layer is configured to apply a first trained model to the first input layer, A second neural network subsystem comprising at least one second input layer and at least one second high-density inner layer, wherein at least one second input layer is configured to receive a first input, and at least one second high-density inner layer is configured to apply a second trained model to the second input layer, A non-transient machine-readable medium according to any one of embodiments 17 to 20, comprising: a third neural network subsystem including at least one third high-density inner layer configured to receive a first output from at least one first high-density inner layer and a second output from at least one second high-density layer, and to apply a third trained model to the first and second outputs to thereby predict a VAR output.
[0136] Embodiment 22. The non-temporary machine-readable medium according to Embodiment 21, wherein at least one first high-density inner layer includes a trained image recognition model and an output high-density inner layer, and at least one second high-density inner layer includes a plurality of second high-density inner layers.
[0137] Embodiment 23. A non-temporal machine-readable medium according to any one of Embodiments 17 to 22, further comprising training a neural network system using two-dimensional imaging data associated with a first group of previously treated subjects and three-dimensional imaging data associated with a second group of previously treated subjects, prior to receiving a first input and a second input.
[0138] Embodiment 24. A non-temporary machine-readable medium according to Embodiment 23, further comprising using visual acuity measurements associated with a second group of subjects who have previously received treatment, demographic statistics associated with a second group of subjects who have previously received treatment, or a combination thereof, to train a neural network system.
[0139] Embodiment 25. A method for predicting a visual acuity response, Receiving input including 2D imaging data associated with the subject undergoing treatment, A method comprising predicting a visual acuity response (VAR) output, which includes a predicted change in the visual acuity of a subject receiving a treatment in response to the treatment, using an input via a neural network system.
[0140] Embodiment 26. The method of Embodiment 25, wherein the two-dimensional imaging data includes color fundus imaging data associated with the subject receiving treatment.
[0141] Embodiment 27. A neural network system, At least one input layer configured to receive input, The method according to embodiment 25 or 26, comprising: at least one high-density inner layer configured to apply a trained model to an input layer, thereby predicting a VAR output.
[0142] Embodiment 28. The method according to Embodiment 27, wherein at least one high-density inner layer includes a trained image recognition model and an output high-density inner layer.
[0143] Embodiment 29. The method according to any one of Embodiments 25 to 28, further comprising training a neural network system using two-dimensional imaging data associated with a plurality of subjects that have previously undergone treatment, prior to receiving input.
[0144] Embodiment 30. A system for predicting visual acuity response, Non-temporary memory and One or more processors coupled to non-temporary memory and configured to read instructions from non-temporary memory to perform an operation on the system, wherein the operation is, Receiving input including 2D imaging data associated with the subject undergoing treatment, A system comprising one or more processors, including predicting a visual acuity response (VAR) output, which includes predicting a change in the visual acuity of a subject receiving treatment in response to the treatment, using inputs via a neural network system.
[0145] Embodiment 31. The system according to Embodiment 30, wherein the two-dimensional imaging data includes color fundus imaging data associated with a subject receiving treatment.
[0146] Embodiment 32. A neural network system, At least one input layer configured to receive input, The system according to embodiment 30 or 31, comprising at least one high-density inner layer configured to apply a trained model to an input layer, thereby predicting a VAR output.
[0147] Embodiment 33. The system according to Embodiment 32, wherein at least one high-density inner layer includes a trained image recognition model and an output high-density inner layer.
[0148] Embodiment 34. The system according to any one of Embodiments 30 to 33, further comprising training a neural network system using two-dimensional imaging data associated with a plurality of subjects that have previously been treated, prior to receiving input.
[0149] Embodiment 35. The system according to Embodiment 34, further comprising using visual acuity measurements associated with a second group of subjects who have previously received treatment, demographic statistics associated with a second group of subjects who have previously received treatment, or a combination thereof, to train the neural network system.
[0150] Embodiment 36. A non-temporary machine-readable medium storing machine-readable instructions that can be executed to cause a system to perform an operation, wherein the operation is Receiving input including 2D imaging data associated with the subject undergoing treatment, A non-temporal machine-readable medium, comprising predicting a visual acuity response (VAR) output, including a predicted change in the visual acuity of a subject receiving a treatment in response to the treatment, using input via a neural network system.
[0151] Embodiment 37. A non-temporary machine-readable medium according to Embodiment 36, wherein the two-dimensional imaging data includes color fundus imaging data associated with a subject receiving treatment.
[0152] Embodiment 38. A neural network system, At least one input layer configured to receive input, A non-transient machine-readable medium according to embodiment 36 or 37, comprising at least one high-density inner layer configured to apply a trained model to an input layer, thereby predicting a VAR output.
[0153] Embodiment 39. A non-transient machine-readable medium according to Embodiment 38, wherein at least one high-density inner layer includes a trained image recognition model and an output high-density inner layer.
[0154] Embodiment 40. A non-temporary machine-readable medium according to any one of Embodiments 36 to 39, further comprising training a neural network system using two-dimensional imaging data associated with a plurality of subjects that have previously been treated, prior to receiving input.
[0155] Embodiment 41. A method for predicting a visual acuity response, Receiving input including 3D imaging data associated with the subject undergoing treatment, A non-temporal machine-readable medium, comprising predicting a visual acuity response (VAR) output, including a predicted change in the visual acuity of a subject receiving a treatment in response to the treatment, using input via a neural network system.
[0156] Embodiment 42. The method according to Embodiment 41, wherein the 3D imaging data includes optical coherence tomography (OCT) imaging data associated with the subject receiving treatment.
[0157] Embodiment 43. The method according to Embodiment 41 or 42, wherein the input further includes visual acuity measurements associated with the subject receiving treatment and demographic statistics associated with the subject receiving treatment.
[0158] Embodiment 44. A neural network system, At least one input layer configured to receive input, The method according to any one embodiment 41 to 43, comprising: at least one high-density inner layer configured to apply a trained model to an input layer, thereby predicting a VAR output.
[0159] Embodiment 45. The method according to Embodiment 44, wherein at least one high-density inner layer comprises multiple high-density inner layers.
[0160] Embodiment 46. The method according to any one of Embodiments 41 to 45, further comprising training a neural network system using three-dimensional imaging data associated with a plurality of subjects who have previously undergone treatment, prior to receiving input.
[0161] Embodiment 47. The method according to Embodiment 46, further comprising using visual acuity measurements associated with a plurality of previously treated subjects, demographic statistics associated with a plurality of previously treated subjects, or a combination thereof, to train a neural network system.
[0162] Embodiment 48. A system for predicting visual acuity response, Non-temporary memory and One or more processors coupled to non-temporary memory and configured to read instructions from non-temporary memory to perform an operation on the system, wherein the operation is, Receiving input including 3D imaging data associated with the subject undergoing treatment, A non-temporal machine-readable medium, comprising predicting a visual acuity response (VAR) output, including a predicted change in the visual acuity of a subject receiving a treatment in response to the treatment, using input via a neural network system.
[0163] Embodiment 49. The system according to Embodiment 48, wherein the 3D imaging data includes optical coherence tomography (OCT) imaging data associated with a subject receiving treatment.
[0164] Embodiment 50. The system according to Embodiment 48 or 49, wherein the input further includes visual acuity measurements associated with the subject receiving treatment and demographic data associated with the subject receiving treatment.
[0165] Embodiment 51. A neural network system, At least one input layer configured to receive input, A system according to any one of embodiments 48 to 50, comprising at least one high-density inner layer configured to apply a trained model to an input layer, thereby predicting a VAR output.
[0166] Embodiment 52. The system according to Embodiment 51, wherein at least one high-density inner layer comprises a plurality of high-density inner layers.
[0167] Embodiment 53. The system according to any one of Embodiments 48 to 52, further comprising training a neural network system using three-dimensional imaging data associated with a plurality of subjects that have previously been treated, prior to receiving input.
[0168] Embodiment 54. The system according to Embodiment 53, further comprising using visual acuity measurements associated with a plurality of previously treated subjects, demographic statistics associated with a plurality of previously treated subjects, or a combination thereof, to train the neural network system.
[0169] Embodiment 55. A non-temporary machine-readable medium storing machine-readable instructions that can be executed to cause a system to perform an operation, wherein the operation is Receiving input including 3D imaging data associated with the subject undergoing treatment, A non-temporal machine-readable medium, comprising predicting a visual acuity response (VAR) output, including a predicted change in the visual acuity of a subject receiving a treatment in response to the treatment, using input via a neural network system.
[0170] Embodiment 56. A non-temporary machine-readable medium according to Embodiment 55, wherein the 3D imaging data includes optical coherence tomography (OCT) imaging data associated with a subject receiving treatment.
[0171] Embodiment 57. A non-temporary machine-readable medium according to Embodiment 55 or 56, wherein the input further includes visual acuity measurements associated with the subject receiving treatment and demographic statistics associated with the subject receiving treatment.
[0172] Embodiment 58. A neural network system, At least one input layer configured to receive input, A non-transient machine-readable medium according to any one of embodiments 55 to 57, comprising at least one high-density inner layer configured to apply a trained model to an input layer, thereby predicting a VAR output.
[0173] Embodiment 59. A non-transient machine-readable medium according to Embodiment 58, wherein at least one high-density inner layer comprises a plurality of high-density inner layers.
[0174] Embodiment 60. A non-temporary machine-readable medium according to any one of Embodiments 55 to 59, further comprising training a neural network system using three-dimensional imaging data associated with a plurality of subjects that have previously been treated, prior to receiving an input.
[0175] Embodiment 61. A non-temporary machine-readable medium according to Embodiment 60, further comprising using visual acuity measurements associated with a plurality of subjects who have previously received treatment, demographic statistics associated with a plurality of subjects who have previously received treatment, or a combination thereof, to train a neural network system.
[0176] Embodiment 62. A method for treating a subject diagnosed with AMD symptoms, Receiving a first input including two-dimensional imaging data associated with the subject, Receiving a second input, which includes 3D imaging data associated with the subject, The method involves generating a treatment output using a first input and a second input via a trained neural network system, wherein the treatment output indicates a predicted change in the subject's visual acuity in response to the treatment. Providing treatment recommendations to healthcare providers based on treatment output, wherein the treatment recommendations are, A method comprising administering a treatment to a subject in response to the treatment output being an improvement in the subject's visual acuity, wherein the step of administering the treatment comprises intravitreal administration of the treatment or a derivative thereof in a treatment dose, wherein the treatment is ranibizumab, and the treatment dose is 0.3 milligrams (mg) or 0.5 mg, and providing a healthcare provider to administer the treatment.
Claims
1. A method for predicting visual acuity response, Receiving a first input which includes two-dimensional imaging data associated with a subject undergoing treatment, Receiving a second input including three-dimensional imaging data associated with the subject undergoing the aforementioned treatment, Predicting visual acuity response (VAR) output, including a predicted change in the visual acuity of the subject undergoing the treatment, using the first and second inputs via a neural network system; Methods that include...
2. The method according to claim 1, wherein the three-dimensional imaging data includes optical coherence tomography (OCT) imaging data associated with the subject receiving the treatment, and the two-dimensional imaging data includes color fundus imaging data associated with the subject receiving the treatment.
3. The method according to claim 1, wherein the second input further includes visual acuity measurements associated with the subject receiving the treatment and demographic statistics associated with the subject receiving the treatment.
4. Predicting the VAR output via the aforementioned neural network system is A first output is generated using the two-dimensional imaging data associated with the subject who has undergone the aforementioned treatment, A second output is generated using the three-dimensional imaging data associated with the subject who has undergone the aforementioned treatment, The VAR output is generated by fusing the first output and the second output, The method according to claim 1, including the method described in claim 1.
5. The aforementioned neural network system A first neural network subsystem comprising at least one first input layer and at least one first high-density inner layer, wherein the at least one first input layer is configured to receive the first input, and the at least one first high-density inner layer is configured to apply a first trained model to the first input layer, A second neural network subsystem comprising at least one second input layer and at least one second high-density inner layer, wherein the at least one second input layer is configured to receive the second input, and the at least one second high-density inner layer is configured to apply a second trained model to the second input layer, A third neural network subsystem comprising at least one third high-density inner layer configured to receive a first output from at least one first high-density inner layer and a second output from at least one second high-density inner layer, and to apply a third trained model to the first and second outputs to predict the VAR output, The method according to claim 1, comprising:
6. The method according to claim 5, wherein the at least one first high-density inner layer comprises a trained image recognition model and an output high-density inner layer, or the at least one second high-density inner layer comprises a plurality of second high-density inner layers.
7. The method according to claim 1, further comprising training the neural network system using two-dimensional imaging data associated with a first group of subjects who have previously received the treatment, and using three-dimensional imaging data associated with a second group of subjects who have previously received the treatment.
8. The method according to any one of claims 1 to 7, wherein the neural network system is to which an attention mechanism is applied.
9. A system for predicting visual acuity responses, Non-temporary memory and One or more processors coupled to the non-temporary memory and configured to read instructions from the non-temporary memory to cause the system to perform the method according to any one of claims 1 to 8, A system that includes these features.
10. A computer program having machine-readable instructions that can be executed to cause a computer to perform the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Apparatus and program for creating retina treatment schedule
JP2013027439A
Method, computer readable memory storing executable program, and apparatus for predicting age-related macular degeneration by image reconstruction
JP2019528826A
Image processing device, image processing method and program
JP2020039851A
Method for measuring the therapeutic effect of a treatment for retinal disease
JP2020511539A
Detection, prediction, and classification for ocular disease
WO2020210891A1