Machine learning enables localization of the foveal center in spectral-domain optical coherence tomography volume scans
A 3D CNN model accurately locates the foveal center in OCT volumes, addressing the limitations of manual and 2D methods by providing precise three-dimensional coordinates, thus improving retinal disease assessment and management.
Patent Information
- Application Number
- JP2025507314
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-08-12
- Filing Date
- 2023-08-14
- Publication Date
- 2025-08-15
AI Technical Summary
Current methods for locating the foveal center in optical coherence tomography (OCT) volumes are inaccurate, time-consuming, and resource-intensive, particularly due to manual corrections and limitations of two-dimensional machine learning models, which can lead to misalignment and reduced accuracy in quantitative measurements.
A three-dimensional convolutional neural network (3D CNN) model is used to automatically detect the foveal center in OCT volumes, providing precise three-dimensional coordinates, reducing computational resources and improving accuracy compared to manual and two-dimensional approaches.
The 3D CNN model enables accurate and reliable localization of the foveal center, enhancing the precision of quantitative measurements and segmentation outputs, leading to improved retinal disease screening, diagnosis, and treatment management.
Smart Images

Figure 2025526684000001_ABST
Abstract
Description
CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application is related to and claims the benefit of the priority date of U.S. Provisional Application No. 63 / 371,297, filed August 12, 2022, entitled "Machine Learning Enabled Localization of Foveal Center in Spectral Domain Optical Coherence Tomography Volume Scans," which is incorporated herein by reference in its entirety. [Technical Field]
[0002] The subject matter described herein relates generally to optical coherence tomography (OCT), and more particularly to an analysis method and system that uses deep learning techniques to detect and locate the center of the fovea in an OCT volume. [Background technology]
[0003] Various imaging techniques have been developed to obtain medical images of tissues, which can then be analyzed to determine the presence or progression of disease. For example, optical coherence tomography (OCT) refers to a technique that uses light waves to obtain two-dimensional slice images (e.g., OCT B-scans) and three-dimensional volume images (e.g., OCT volumes) of tissues, such as a patient's retina. OCT images can be analyzed to calculate quantitative measurements for use in screening, diagnosing, and managing treatment for retinal disease. Such quantitative measurements can include, for example, measurements related to the center of the fovea of the retina, such as central subfield thickness (CST) and other retinal thickness measurements (e.g., thickness measurements relative to the Early Treatment Diabetic Retinopathy Study (ETDRS) grid). Therefore, it would be desirable to have a method and system that improves the accuracy and reliability of quantitative measurements calculated based on the center of the fovea of the retina using OCT images. Summary of the Invention
[0004] In one or more embodiments, a method is provided in which an optical coherence tomography (OCT) volume of a subject's retina is received. The OCT volume includes multiple OCT B-scans of the retina. A three-dimensional image input of a model is generated using the OCT volume. The model includes a three-dimensional convolutional neural network. The model is used to generate a foveal center location, including three-dimensional coordinates of the center of the fovea of the retina, based on the three-dimensional image input.
[0005] In one or more embodiments, a method for training a model is provided. The method includes receiving a training dataset including a plurality of optical coherence tomography (OCT) volumes for a plurality of retinas, each of the plurality of OCT volumes including a plurality of OCT B-scans. A three-dimensional training image input is generated for the model using the plurality of OCT volumes of the training dataset, the model including a three-dimensional convolutional neural network and a recurrent layer. The model is trained to generate a foveal center location including three-dimensional coordinates of the center of the fovea of the retina of the selected OCT volume based on the training three-dimensional image input.
[0006] In one or more embodiments, a system is provided that includes one or more data processors and a non-transitory computer-readable storage medium that includes instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of one or more methods disclosed herein. For example, the one or more data processors may be caused to receive an optical coherence tomography (OCT) volume of a subject's retina, the OCT volume including multiple OCT B-scans of the retina, generate a three-dimensional image input of a model using the OCT volume, the model including a three-dimensional convolutional neural network, and generate, via the model, a foveal center location including three-dimensional coordinates of the center of the fovea of the retina based on the three-dimensional image input.
[0007] In one or more embodiments, a computer program product is provided that is tangibly embodied in a non-transitory machine-readable storage medium and includes instructions configured to cause one or more data processors to perform some or all of one or more of the methods disclosed herein. For example, the one or more data processors may be caused to receive an optical coherence tomography (OCT) volume of a subject's retina, the OCT volume including multiple OCT B-scans of the retina, generate a three-dimensional image input of a model using the OCT volume, the model including a three-dimensional convolutional neural network, and generate, via the model, a foveal center location including three-dimensional coordinates of the center of the fovea of the retina based on the three-dimensional image input.
[0008] The terms and expressions which have been employed are used as terms of description rather than of limitation, and there is no intention in the use of such terms and expressions to exclude equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the invention as claimed. Thus, although the claimed invention has been specifically disclosed by embodiments and optional features, it will be understood that modifications and variations of the concepts disclosed herein may be resorted to by those skilled in the art, and that such modifications and variations are deemed to be within the scope of the invention as defined by the appended claims. [Brief explanation of the drawings]
[0009] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate certain aspects of the subject matter disclosed herein and, together with the description, serve to explain some of the principles associated with the disclosed embodiments.
[0010] [Figure 1] FIG. 1 is a block diagram of an image processing system in accordance with one or more exemplary embodiments.
[0011] [Figure 2] 1 is a flowchart of a process for processing an optical coherence tomography (OCT) volumetric image of a subject's retina, according to one or more exemplary embodiments.
[0012] [Figure 3] 1 is a flowchart of a process for training a model to generate a foveal center location, according to one or more embodiments.
[0013] [Figure 4] FIG. 1 illustrates an exemplary workflow for processing an OCT volume, according to one or more exemplary embodiments.
[0014] [Figure 5] FIG. 1 is a block diagram illustrating an example of a computing system, in accordance with one or more exemplary embodiments.
[0015] It should be understood that the figures are not necessarily drawn to scale, and that objects in the figures are not necessarily drawn to scale relative to each other. The figures are intended to provide clarity and understanding of various embodiments of the devices, systems, and methods disclosed herein. Wherever possible, the same reference numerals will be used throughout the figures to refer to the same or like parts. Furthermore, it should be appreciated that the figures are not intended to limit the scope of the present teachings in any way. DETAILED DESCRIPTION OF THE INVENTION
[0016] I. Overview Medical imaging technologies are powerful tools that can be used to generate medical images that allow medical professionals to better visualize and understand a patient's medical problems, thereby providing more accurate diagnostic and treatment options. To accurately diagnose and treat retinal diseases and more generally monitor the retina, these medical imaging technologies can be used to image specific portions of the retina, such as the fovea.
[0017] The fovea (or fovea centralis) is a small depression in the center of the macula of the retina. It is formed by a densely packed set of cones and is surrounded by a parafoveal belt, a perifoveal outer region, and a larger peripheral region. The center of the fovea (or foveal center) is occupied by the highest density of cones found in the retina, whereas cone density is significantly lower in the fovea. For example, the cone density in the perifoveal outer region is approximately 12 cones per 100 micrometers, whereas the most central part of the fovea is approximately 50 cones per 100 micrometers. Although a relatively small portion of the retina, the fovea is responsible for acute central vision (e.g., foveal vision), which is essential for activities that rely on visual detail. For example, orienting the fovea in a particular direction can focus sensory processing resources on the most relevant information sources.
[0018] The center of the fovea (or foveal center) is an important retinal feature for understanding disease states and vision loss. For example, as noted, the center of the fovea is densely packed and contains the highest density of cones found in the retina. Retinal thickening, including swelling at or around the center of the fovea, is typically associated with vision loss.
[0019] Therefore, the center of the fovea (or foveal center) is an important landmark for generating further analyses of retinal characteristics. Identifying the location of the foveal center can be essential for measuring biomarkers related to diagnosing retinal disease, assessing disease burden, monitoring retinal disease progression, and predicting treatment response. For example, central subfield thickness (CST), an important quantitative measure for disease monitoring and treatment response, can be determined based on the foveal center of the retina. CST is typically measured by calculating the average retinal thickness over a circular area (e.g., a 1-millimeter ring) centered around the foveal center. Similarly, measurements related to retinal segmentation can be based on the foveal center, such as, but not limited to, generating an Early Treatment Diabetic Retinopathy Study (ETDRS) grid or calculating the boundary segmentation of the internal limiting membrane (ILM) and / or Bruch's membrane (BM).
[0020] Optical coherence tomography (OCT) is a noninvasive imaging technique that is particularly popular for obtaining images of the retina. OCT is sometimes described as an ultrasound scanning technique that scatters light waves from tissue to produce OCT images in the form of two-dimensional (2D) and / or three-dimensional (3D) images of tissue, similar to ultrasound scanning, which uses sound waves to scan tissue. 2D OCT images are sometimes referred to as OCT slices, OCT cross-sectional images, or OCT scans (e.g., OCT B-scans). 3D OCT images are sometimes referred to as OCT volume images and may be composed of many OCT slice images.
[0021] For example, the presence of a patient's retinal disease, such as neovascular age-related macular degeneration, diabetic macular edema, or some other type of retinal disease, can cause a misalignment between the center of the patient's fovea and the geometric center of the resulting OCT volume. Similarly, poor subject fixation, eye movement, head tilt, subject age, or a combination thereof can shift the center of the subject's fovea relative to the geometric center of the resulting OCT volume. A fixation error occurs when the subject's fixation on the target does not overlap with the center of the fovea. Such fixation errors can be due to, for example, neurological conditions, ocular tilt, subject age, etc. In some cases, a misalignment of the foveal center can exist due to a lack of training or experience on the part of the clinician performing the OCT scan. For example, medical students at a teaching hospital may be less successful at aligning the center of the subject's fovea with the geometric center of the OCT scan compared to experienced OCT clinicians.
[0022] Currently, human raters perform manual correction of the foveal center to prevent further problems when using the foveal center to perform quantitative measurements and identify biomarkers. However, manual correction can be time-consuming, undesirable, or even infeasible for large datasets. Furthermore, manual correction may not have the desired level of accuracy. For example, significant inter- and intra-clinician variability can exist in subsequent recalibration efforts. Furthermore, these manual-based approaches to locating the foveal center rely on expert annotation of optical coherence tomography (OCT) B-scans, which can be resource-intensive and error-prone.
[0023] When an expert or manual rater attempts to detect the center of the fovea, the rater isolates and identifies the OCT B-scan of the OCT volume that is most likely to display the foveal center, which is typically the mid- or central OCT B-scan relative to the transverse axis (e.g., each OCT B-scan in the OCT volume may be at a different location along the transverse axis). From that selected OCT B-scan, the rater manually locates the fovea on the OCT B-scan. For example, the rater may mark the transverse location of the foveal center and then infer the axial location of the foveal center. However, due to misalignment, the manually identified OCT B-scan may not be the correct OCT B-scan containing the foveal center. The rater must then review multiple OCT B-scans surrounding the central OCT B-scan (relative to the transverse axis) to find the foveal center. Attempting to locate the foveal center in a large dataset containing hundreds or thousands of OCT volumes can be challenging. Furthermore, in OCT images acquired from the retina, the presence of lesions (or abnormalities) can alter the expected appearance of the fovea, making it difficult to detect the fovea, even for experienced human raters.
[0024] Some currently available methods for locating the foveal center use two-dimensional machine learning models (e.g., two-dimensional convolutional neural networks) to automatically detect the foveal center. Such models are trained using a single OCT B-scan of an OCT volume for iteration. This OCT B-scan is either the geometric center of the OCT volume or an OCT B-scan selected by a human evaluator to correct for misalignment. However, as discussed above, this correction may not be accurate. Therefore, training a model based on such an OCT B-scan may result in the model's accuracy being reduced in detecting the foveal center. Furthermore, these currently existing two-dimensional machine learning models use pixel-by-pixel classification to detect the foveal center. Each pixel in the input OCT B-scan is assigned a classification (or probability) of whether the pixel is likely to be the foveal center or not. Therefore, post-processing resources are required for these models to convert this pixel-by-pixel classification of the various OCT B-scans of the OCT volume into simple coordinates of the foveal center. Therefore, these types of techniques may be undesirable in many scenarios.
[0025] Therefore, the embodiments described herein recognize the importance of having improved methods and systems for automatically detecting and locating (e.g., determining coordinates of) the foveal center in an OCT volume accurately, precisely, and reliably. The embodiments described herein provide methods and systems for automatically detecting and generating three-dimensional coordinates of the foveal center accurately, precisely, and reliably. In one or more embodiments, an optical coherence tomography (OCT) volume of a subject's retina is received. The OCT volume includes multiple OCT B-scans of the retina. The OCT volume is used to generate a three-dimensional image input of a model. The model includes a three-dimensional convolutional neural network and a recurrent layer. The recurrent layer may be considered part of the three-dimensional convolutional neural network or may be separate. The model is used to generate a foveal center location. The foveal center location includes three-dimensional coordinates of the center of the fovea of the retina based on the three-dimensional image input. In other embodiments, the foveal center location includes two coordinates, such as a transverse coordinate and a lateral coordinate. This foveal center location of the foveal center is generated with a higher level of accuracy than a manually identified foveal center. Furthermore, this foveal center location may be more accurate and reliable than the foveal center identified based on pixel-by-pixel classification of OCT B-scans.
[0026] The foveal center location generated by the model can be used to generate several outputs. For example, quantitative measurements can be automatically calculated from the OCT volume to be more accurate and reliable. Such quantitative measurements include, but are not limited to, central subfield thickness (CST) measurements and various retinal thickness measurements calculated using a retinal grid. The retinal grid can be, for example, but is not limited to, an ETDRS grid.
[0027] In some embodiments, a model system is used that includes various modules or layers for automatically identifying the location of the foveal center in two or three dimensions and further automatically identifying several quantitative measurements (e.g., CST measurements, one or more retinal thickness measurements of a retinal grid, etc.). The model system can also be used to automatically improve segmentation outputs, such as segmentation of various retinal layers or retinal features in an OCT volume (or in various OCT B-scans of an OCT volume). These segmentation outputs can have improved accuracy due to the more accurately identified location of the foveal center.
[0028] Thus, the embodiments described herein enable accurate and reliable localization of the foveal center in at least two dimensions. Because the embodiments described herein use a model to automatically and directly calculate and output the coordinates of the foveal center of an OCT volume without requiring further processing, the amount of computational resources used may be reduced and the overall process may be less laborious. Locating the foveal center further enables the generation of more accurate and reliable quantitative measurement and segmentation outputs, which in turn may lead to improved retinal disease screening, diagnosis, management, treatment selection, treatment response prediction, treatment management, or a combination thereof.
[0029] II. Exemplary System for Foveal Center Detection 1 is a block diagram of an image processing system 100 according to one or more exemplary embodiments. Image processing system 100 may be used to process ophthalmic images to extract features from such images, correct or adjust one or more features extracted from such images, segment such images, and generate one or more outputs related to the diagnosis, screening, and / or treatment of ophthalmic disorders, or a combination thereof.
[0030] Image processing system 100 includes analysis system 101. Analysis system 101 may be implemented using hardware, software, firmware, or a combination thereof. In one or more embodiments, analysis system 101 may include computing platform 102, data storage 104 (e.g., a database, a server, a storage module, cloud storage, etc.), and display system 106. Computing platform 102 may take various forms. In one or more embodiments, computing platform 102 includes a single computer (or computer system) or multiple computers that communicate with each other. In other examples, computing platform 102 takes the form of a cloud computing platform, a mobile computing platform (e.g., a laptop, a smartphone, a tablet, etc.), another processor-based device (e.g., a workstation or desktop computer), a wearable computing device (e.g., a smart watch), etc., or a combination thereof.
[0031] Data storage 104 and display system 106 each communicate with computing platform 102. In some examples, data storage 104, display system 106, or both may be considered part of computing platform 102 or may be otherwise integrated with computing platform 102. Thus, in some examples, computing platform 102, data storage 104, and display system 106 may be separate components that communicate with each other, while in other examples, some combination of these components may be integrated together.
[0032] The computing platform 102 may be or may be part of a client device, which may be, for example, a processor-based device including a workstation, a desktop computer, a laptop computer, a smartphone, a tablet computer, a wearable device, etc.
[0033] The image processing system 100 may further include an OCT imaging system 110, which may also be referred to as an OCT scanner. The OCT imaging system 110 may generate spectral domain (SD) OCT images. The OCT imaging system 110 may generate OCT imaging data 108. The OCT imaging data 108 may include any number of three-dimensional, two-dimensional, or one-dimensional OCT images. A three-dimensional OCT image may be referred to as an OCT volume. A two-dimensional OCT image may take the form of, for example, but not limited to, an OCT B-scan.
[0034] In one or more embodiments, the OCT imaging data 108 includes an OCT volume 114 (e.g., an SD-OCT volume) of the subject's retina. The OCT volume 114 may be composed of multiple OCT B-scans 115 of the subject's retina. The multiple OCT B-scans 115 may include, for example, without limitation, 10s, 100s, 1000s, 10,000s, or some other number of OCT B-scans. The OCT B-scans may also be referred to as OCT slice images or cross-sectional OCT images.
[0035] In some embodiments, the retina is a healthy retina. In other embodiments, the retina has been diagnosed with a retinal disease. For example, the diagnosis may be one of age-related macular degeneration (AMD), neovascular age-related macular degeneration (nAMD), diabetic retinopathy, macular edema, geographic atrophy, or any other type of retinal disease.
[0036] In one or more embodiments, the OCT imaging system 110 includes an optical coherence tomography (OCT) system (e.g., an OCT scanner or machine) configured to generate OCT imaging data 108 of a patient's tissue. For example, the OCT imaging system 110 may be used to generate OCT imaging data 108 of a patient's retina. In some cases, the OCT imaging system 110 may be a large tabletop configuration used in a clinical setting, a portable or handheld dedicated system, or a "smart" OCT system integrated into a user's personal device such as a smartphone. In some cases, the OCT imaging system 110 may include an image denoiser configured to remove noise and other artifacts from the raw OCT volumetric images to generate the OCT volume 114.
[0037] Analysis system 101 may communicate with OCT imaging system 110 via network 112. Network 112 may be implemented using a single network or a combination of multiple networks. Network 112 may be implemented using any number of wired, wireless, or optical communication links, or a combination thereof. For example, in various embodiments, network 112 may include the Internet or one or more intranets, landline networks, wireless networks, and / or other suitable types of networks. In another example, network 112 may comprise a wireless telecommunications network (e.g., a cellular network) adapted to communicate with other communication networks, such as the Internet. In some cases, network 112 includes at least one of a local area network (LAN), a virtual local area network (VLAN), a wide area network (WAN), a public land mobile network (PLMN), the Internet, or another type of network.
[0038] OCT imaging system 110 and analysis system 101 may each include one or more electronic processors, electronic memory, and other suitable electronic components for executing instructions, such as program code and / or data stored on one or more computer-readable media, to implement the various applications, data, and steps described herein. For example, such instructions may be stored on one or more computer-readable media, such as internal and / or external memory or data storage devices (e.g., data storage 104) of the various components of image processing system 100, and / or may be accessible via network 112.
[0039] While only one of each OCT imaging system 110 and analysis system 101 is shown, in other embodiments, there may be two or more of each. Furthermore, while FIG. 1 depicts OCT imaging system 110 and analysis system 101 as two separate components, in some embodiments, OCT imaging system 110 and analysis system 101 may be part of the same system (e.g., maintained by the same entity, such as a healthcare provider or clinical trial administrator). In some cases, portions of analysis system 101 may be implemented as part of OCT imaging system 110. For example, analysis system 101 may be configured to operate as a module implemented using a processor, microprocessor, or some other hardware component of OCT imaging system 110. In still other embodiments, analysis system 101 may be implemented within a cloud computing system that can be accessed by or otherwise communicate with OCT imaging system 110.
[0040] Analysis system 101 may include image processor 116 configured to receive OCT imaging data 108 from OCT imaging system 110. Image processor 116 may be implemented using hardware, firmware, software, or a combination thereof. In one or more embodiments, image processor 116 may be implemented within computing platform 102. In some cases, at least a portion of image processor 116 (e.g., its modules) is implemented within OCT imaging system 110.
[0041] In one or more embodiments, the image processor 116 may generate a three-dimensional (3D) image input 118 using the OCT imaging data 108. For example, the OCT volume 114 may be preprocessed using a set of preprocessing operations to form the 3D image input 118. The set of preprocessing operations may include, for example, but not limited to, at least one of a normalization operation, a scaling operation, a resizing operation, a horizontal flip operation, a vertical flip operation, a cropping operation, a rotation operation, a noise filtering operation, or some other type of preprocessing operation. A normalization operation may be performed to normalize the coordinates of the coordinate system of the OCT volume 114. In some cases, pixel values may be normalized (e.g., normalized to values between 0 and 1). The scaling operation may include, for example, scaling the coordinate system associated with the OCT volume 114. The resizing operation may include resizing each of the multiple OCT B-scans 115. A pre-processing operation of the set of pre-processing operations may be performed on one or more of the OCT B-scans 115 of the OCT volume 114 .
[0042] The image processor 116 may include a model 120, which may be referred to as a foveal center model. The model 120 is a deep learning model that includes a 3D convolutional neural network (CNN) 122 and a set of output layers 124. The 3D CNN 122 may include a fully convolutional neural network. The set of output layers 124 includes at least one recurrent layer. For example, the 3D CNN 122 may take a 3D image input 118 as input and output spatial information. This spatial information may take the form of an output image indicating the location of the foveal center. The recurrent layer is used to convert this output image into n-dimensional coordinates. The recurrent layer may be considered part of the 3D convolutional neural network or may be separate.
[0043] In one or more embodiments, the model 120 is a lightweight model. For example, the 3D CNN 122 may be a lightweight convolutional neural network. A lightweight convolutional neural network has reduced computational complexity. This reduction can be performed in various ways, including, for example, reducing or eliminating dense layers.
[0044] The model 120 is used to detect and generate a foveal center location 126 based on the 3D image input 118. The foveal center location 122 includes a three-dimensional coordinate 128 (i.e., three coordinates for three different axes of a coordinate system) relative to the center of the fovea of the retina. In one or more embodiments, the three-dimensional coordinates 128 may be relative to a selected coordinate system corresponding to the OCT volume 114. For example, the three-dimensional coordinates 128 may include a transverse coordinate for a transverse axis, a lateral coordinate for a lateral axis, and an axial coordinate for an axial axis. The transverse axis may be the axis along which each of the multiple OCT B-scans 115 lies. The lateral axis and the axial axis may be the axes of each pixel of the multiple OCT B-scans 115. In some cases, each of the multiple OCT B-scans may be indexed by the transverse axis (e.g., may have an index corresponding to its value). The lateral axis may be, for example, the horizontal axis of each OCT B-scan. The axial axis may be the vertical axis of each OCT B-scan. In other embodiments, the foveal center location 126 includes two coordinates for the center of the fovea. For example, the foveal center location 126 may include a transverse coordinate and a lateral coordinate of the center of the fovea.
[0045] In some embodiments, model 120 may include a layer or module for rounding a row coordinate of three-dimensional coordinates 128 of foveal center location 126 to a value corresponding to an index associated with a particular OCT B-scan of multiple OCT B-scans 115. For example, an initial value of the row coordinate of foveal center location 126 may be rounded to a rounded value corresponding to an index associated with a particular OCT B-scan of multiple OCT B-scans 115. This rounded value forms one of three-dimensional coordinates 128 of foveal center location 126.
[0046] In one or more embodiments, the model 120 may be trained with a training dataset 142 to generate the foveal center location 126. The training dataset 142 includes a plurality of OCT volumes. One or more OCT volumes in the training dataset 142 may capture a healthy retina. One or more OCT volumes in the training dataset 142 may capture a retina diagnosed with a retinal disease (e.g., AMD, nAMD, diabetic retinopathy, macular edema, geographic atrophy, or some other type of retinal disease). An example of a method for training the model 120 (e.g., training the 3D CNN 122) may be further described below in FIG. 3.
[0047] Output generator 130 receives and processes foveal center location 126 to generate output 132 based on foveal center location 126. Output 132 may take a variety of forms. For example, output 132 may include at least one of a corrected foveal center location 134, a central subfield thickness (CST) measurement 136, a retinal grid 138, a report 140 identifying any one or more of foveal center location 126, corrected foveal center location 134, a central subfield thickness (CST) measurement 136, a retinal grid 138, or a combination thereof.
[0048] The output generator 130 may process the foveal center location 126 to generate a corrected foveal center location 134, which may include, for example, the coordinate of the center of the fovea adjusted after transformation. For example, the three-dimensional coordinates 128 may be transformed from a first selected coordinate system associated with the OCT volume 114 to a second selected coordinate system. The second selected coordinate system may be associated with the retina or the object (e.g., an anatomical coordinate system), for example. In some cases, the corrected foveal center location 134 may be generated by rounding one or more of the three-dimensional coordinates 128 of the foveal center location 126 to an integer or a desired decimal level. In some cases, a row coordinate of the three-dimensional coordinates 128 is rounded to a value corresponding to an index associated with a particular OCT B-scan of the plurality of OCT B-scans 115 to form the corrected foveal center location 134.
[0049] The output generator 130 may process the foveal center location 126 to generate a CST measurement 136, which may be a measurement or indication of foveal thickness. The central subfield is a circular region 1 mm in diameter centered at the center of the fovea. The CST measurement 136 is a measurement of macular thickness in the central subfield. This thickness may be an average or mean thickness (e.g., for a selected number of OCT B-scans around the center of the fovea). CST may also be referred to as central macular thickness or mean macular thickness.
[0050] The output generator 130 may process the foveal center location 126 to determine (or otherwise define) a retinal grid 138. The retinal grid 138 is a grid that divides the retina into regions based on the three-dimensional coordinates 128 of the foveal center location 126. In one or more embodiments, the retinal grid 138 takes the form of a retinal thickness grid. Such a retinal thickness grid enables quantitative measurements that may be important for screening for retinal disease, diagnosing retinal disease, monitoring disease progression, monitoring treatment response, selecting a treatment or treatment protocol, or a combination thereof. Such a retinal thickness grid enables these quantitative measurements to function as biomarkers.
[0051] As an example, the retinal grid 138 may be an Early Treatment Diabetic Retinopathy Study (ETDRS) grid. This grid divides the retina into nine regions (a central region, a set of four central regions, and the centers of four outer regions) centered on the center of the fovea. Specifically, the central region is defined as the volume within a 1 mm diameter from the center of the fovea. The set of four central regions is defined as four quadrants within the volume between the central region and a 3 mm diameter perimeter around the center of the fovea. The set of four outer regions is defined as four quadrants within the volume between the set of central regions and a 6 mm diameter perimeter around the center of the fovea. The output generator 130 may use the accurately generated foveal center location 126 to accurately define the dimensions of the ETDRS grid. Quantitative measurements generated using an accurately defined ETDRS grid may improve capabilities for screening, diagnosis, monitoring disease progression, treatment selection, treatment response prediction, treatment management, or a combination thereof.
[0052] In one or more embodiments, the output generator 130 may generate a report 140 that includes any of the above-identified outputs and / or other information. For example, the report 140 may include a reproduced or modified version of a particular OCT B-scan that includes the foveal center, along with a graphical annotation or label indicating the foveal center location 126. For example, the report 140 may modify the OCT B-scan by at least one of resizing, flipping (horizontally, vertically, or both), cropping, rotating, reducing noise, adding graphical features (e.g., adding one or more labels, colors, text, etc.), or otherwise modifying the OCT B-scan. In some embodiments, the foveal center location 126 and / or the modified foveal center location 134 may be identified in the reproduced or modified OCT B-scan via text identifying the corresponding coordinates. In other cases, the foveal center location 126 and / or the modified foveal center location 134 may be identified by a pointer, dot, marker, or some other type of graphical feature.
[0053] In one or more embodiments, analysis system 101 stores OCT imaging data 108 obtained from OCT imaging system 110, 3D image input 118, foveal center location 126, output 132, other data generated during processing of OCT imaging data 108, or a combination thereof, in data storage 104. In some embodiments, the portion of data storage 104 that stores such information may be configured to comply with Health Insurance Portability and Accountability (HIPAA) security requirements, which mandate specific security procedures when handling patient data (e.g., OCT images of patient tissue, etc.) (i.e., data storage 104 may be HIPAA-compliant). For example, the stored information may be encrypted and anonymized. For example, OCT volume 114 may be encrypted and processed to remove and / or obfuscate personal identifying information of the subject from whom OCT volume 114 was obtained. In some cases, the communication link between OCT imaging system 110 and analysis system 101 utilizing network 112 may also be HIPAA-compliant. For example, at least a portion of network 112 may be a virtual private network (VPN) that is end-to-end encrypted and configured to anonymize personally identifiable information data transmitted thereover.
[0054] Image processing system 100 may include any number or combination of servers and / or software components that operate to perform various processes associated with acquiring and processing retinal OCT volumes. Examples of servers may include, for example, standalone and enterprise-class servers. In one or more embodiments, image processing system 100 may be operated and / or maintained by one or more different entities.
[0055] In some embodiments, OCT imaging system 110 may be maintained by an entity tasked with obtaining OCT imaging data 108 of a subject's tissue sample for purposes of disease screening, diagnosis, disease monitoring, disease treatment, research, clinical trial management, or a combination thereof. For example, the entity may be a healthcare provider (e.g., an eye care provider) seeking to obtain OCT imaging data 108 about a subject's retina for use in diagnosing retinal diseases and / or other types of eye conditions. As another example, the entity may be a clinical trial administrator responsible for collecting OCT imaging data 108 about a subject's retina to monitor retinal changes over the course of a disease, monitor treatment response, or both. Analysis system 101 may be maintained by the same or a different entity (or multiple entities) as OCT imaging system 110. For example, analysis system 101 may be maintained by an entity tasked with identifying or discovering retinal disease biomarkers from OCT images.
[0056] III. Exemplary Methodology for Localizing the Foveal Center FIG. 2 is a flowchart of a process for processing an OCT volume of a subject's retina, according to one or more exemplary embodiments. Process 200 of FIG. 2 may be implemented using analysis system 101 of FIG. 1. In one or more embodiments, at least some of the steps of process 200 may be performed by a processor of a computer or server implemented as part of analysis system 101. It will be understood that additional steps may be performed before, between, or after the steps of process 200 described below. Furthermore, in some embodiments, one or more of the steps may be omitted or performed in a different order.
[0057] Process 200 may optionally include step 201 of training a model including a 3D convolutional neural network (CNN). The model may be, for example, model 120 of Figure 1. The 3D CNN may be implemented, for example, using 3D CNN 122 of Figure 1.
[0058] Step 202 of process 200 includes receiving an optical coherence tomography (OCT) volume of a subject's retina, the OCT volume including multiple OCT B-scans of the retina. The OCT volume may be, for example, OCT volume 114 of FIG. 1. The multiple OCT B-scans may be, for example, multiple OCT B-scans 115 of FIG. 1. Each of the multiple OCT B-scans is a cross-sectional view of the retina taken at a particular location relative to a selected axis, sometimes referred to as the transverse axis. Each OCT B-scan may have a horizontal axis (lateral axis) and a vertical axis (axial axis).
[0059] The retina can be a healthy retina. Alternatively, the retina has been diagnosed with retinal disease or is suspected to have retinal disease. The retinal disease can be, for example, age-related macular degeneration (AMD), neovascular age-related macular degeneration (nAMD), diabetic retinopathy, macular edema, geographic atrophy, or any other type of retinal disease.
[0060] Step 204 of process 200 includes generating a 3D image input of the model using the OCT volume. The 3D image input can be, for example, 3D image input 118 of FIG. 1 . The segmented model can be, for example, segmented model 120 of FIG. 1 . The model includes a 3D convolutional neural network and a recurrent layer at the output of the 3D convolutional neural network. The recurrent layer can be considered part of the 3D convolutional neural network or can be separate. Step 204 can be performed in various ways. In one or more embodiments, generating the 3D image input includes performing a set of preprocessing operations on the OCT volume. The set of preprocessing operations can include, for example, at least one of a normalization operation, a scaling operation, a resizing operation, a horizontal flip operation, a vertical flip operation, a cropping operation, a rotation operation, a noise filtering operation, or some other type of preprocessing operation.
[0061] Step 206 includes generating, via the model, a foveal center location including three-dimensional coordinates of the center of the retina's fovea based on the three-dimensional image input. The foveal center location may be, for example, foveal center location 126 comprised of three-dimensional coordinates 128 of FIG. 1 . In one or more embodiments, the three-dimensional coordinates of the foveal center location include a transverse coordinate, an axial coordinate, and a lateral coordinate relative to a selected coordinate system of the OCT volume. In some cases, the model generates three-dimensional coordinates having values to at least a selected decimal level. One or more of the three-dimensional coordinates may be rounded up or down to the selected decimal level. In some cases, one or more of the three-dimensional coordinates are rounded up or down to the nearest integer. In some embodiments, the initial value of the transverse coordinate of the foveal center location is rounded to a rounded value corresponding to an index associated with a particular OCT B-scan among the multiple OCT B-scans. The rounded value becomes the transverse coordinate of the foveal center location.
[0062] Process 200 may optionally include step 208. Step 208 includes generating an output using the predicted foveal center location. Step 208 may be performed using one or more substeps performed simultaneously, sequentially, or in some other type of combination. The output may be, for example, output 132 of FIG. 1. In one or more embodiments, the output includes a central subfield thickness measurement generated using the three-dimensional coordinates of the foveal center location. In some embodiments, the output includes a retinal grid that divides the retina into regions based on the three-dimensional coordinates of the foveal center location. The retinal grid may be, for example, an ETDRS grid that divides the retina into nine regions centered about the three-dimensional coordinates of the foveal center location.
[0063] In some embodiments, the output includes a corrected foveal center location. For example, the foveal center location generated by the model in step 206 may be corrected to form the corrected foveal center location. In some cases, the three-dimensional coordinates of the foveal center location are transformed from a first selected coordinate system associated with the OCT volume to a second selected coordinate system associated with the retina or the object (e.g., an anatomical coordinate system). In some cases, one or more of the three-dimensional coordinates are rounded up or down to a selected decimal level. In some cases, one or more of the three-dimensional coordinates are rounded up or down to the nearest integer. In an exemplary embodiment, the row coordinate of the three-dimensional coordinate of the foveal center location is rounded to a value corresponding to an index associated with a particular OCT B-scan among multiple OCT B-scans of the retina.
[0064] In one or more embodiments, the output includes the transformation necessary to shift the geometric center of the OCT volume to the foveal center position. For example, the transformation can be a three-dimensional (or two-dimensional) shift that can be applied to the geometric center of the OCT volume to arrive at the foveal center position. This transformation can be applied to the foveal center identified by the OCT imaging system (e.g., OCT imaging system 110), which is at the geometric center of the OCT volume, so that the new foveal center is accurate when used for further analysis or processing.
[0065] In one or more embodiments, the output generated in optional step 208 includes a modified segmentation of the OCT volume. For example, the foveal center location may be used to modify the segmentation of at least one OCT B-scan of the multiple OCT B-scans of the retina based on the three-dimensional coordinates of the foveal center location. In some cases, the segmented image output (e.g., a mask image output), which may be two-dimensional or three-dimensional, is modified based on the foveal center location.
[0066] In one or more embodiments, the output includes a report including any one or more of the outputs described above. In some cases, the report may include a reproduced or modified version of a particular OCT B-scan including the foveal center, along with a graphical annotation or label indicating the foveal center location. For example, the report may modify the OCT B-scan by at least one of resizing, flipping (horizontally, vertically, or both), cropping, rotating, reducing noise, adding graphical features (e.g., adding one or more labels, colors, text, etc.), or otherwise modifying the OCT B-scan. In some embodiments, the foveal center location and / or the modified foveal center location may be identified in the reproduced or modified OCT B-scan via text identifying the corresponding coordinates. In other cases, the foveal center location and / or the modified foveal center location may be identified by a pointer, dot, marker, or some other type of graphical feature.
[0067] Process 200, which may be implemented using image processing system 100 described in FIG. 1 or at least analysis system 101 of FIG. 1, provides improvements to the art of retinal disease screening, diagnosis, and treatment management. For example, by improving the accuracy, precision, and reliability of locating the foveal center, process 200 thereby improves the accuracy, precision, and reliability of quantitative measurements made based on foveal center location, improves the identification of biomarkers for retinal disease and / or treatment response, and improves segmentation of OCT images. These improvements may be achieved regardless of the type of OCT imaging system used to generate the OCT volume, the type of scanning protocol used to generate the OCT volume, the quality of the OCT volume, or the type of lesion or abnormality present in the retina. Thus, process 200 may facilitate improved automated analysis of large datasets of OCT volumes, even in the presence of various conditions associated with the OCT volumes.
[0068] FIG. 3 is a flowchart of a process for training a model to generate a foveal center location, according to one or more embodiments. Process 300 of FIG. 3 may be implemented using analysis system 101 of FIG. 1. Process 300 may be an example of an implementation of step 201 of process 200 of FIG. 2. Furthermore, it is understood that additional steps may be performed before, between, or after the steps of process 300 described below. Furthermore, in some embodiments, one or more of the steps may be omitted or performed in a different order.
[0069] Step 302 of process 300 includes receiving a training dataset including multiple optical coherence tomography (OCT) volumes for multiple retinas. Each of the multiple OCT volumes includes multiple OCT B-scans. The training dataset may be, for example, training dataset 142 of FIG. 1. The training dataset may include OCT volumes of retinas in various health states. In one or more embodiments, the training dataset may include one or more OCT volumes for healthy retinas. In one or more embodiments, the training dataset may include OCT volumes of retinas diagnosed with retinal diseases such as AMD, nAMD, diabetic retinopathy, macular edema, geographic atrophy, or some other type of retinal disease. In some cases, the training data may include one or more OCT volumes for damaged retinas. The training dataset may include OCT volumes of the same type of retina (e.g., healthy, diseased, or damaged) or different types of retinas. In some cases, the multiple OCT volumes may be generated by two or more OCT imaging systems (or types of OCT imaging systems).
[0070] Step 304 includes generating a training 3D image input for the model using a plurality of OCT volumes of the training dataset, the model including a 3D convolutional neural network and a recurrent layer. The recurrent layer may be considered part of the 3D convolutional neural network or may be separate. For example, step 304 may include performing a set of preprocessing operations on the plurality of OCT volumes to form the training 3D image input. The set of preprocessing operations may include at least one of a normalization operation, a scaling operation, a resizing operation, a horizontal flip operation, a vertical flip operation, a cropping operation, a rotation operation, a noise filtering operation, or some other type of preprocessing operation.
[0071] Step 306 includes training the model to generate a foveal center location, including the three-dimensional coordinates of the center of the retinal fovea of the selected OCT volume based on the training three-dimensional image input. The foveal center location may include a transverse coordinate, a lateral coordinate, and an axial coordinate of the center of the fovea of the OCT volume. The transverse coordinate may correspond to a specific OCT B-scan of the selected OCT volume. In some cases, the transverse coordinate may be a rounded value that corresponds to a specific OCT B-scan. In other cases, the transverse coordinate may be a value that can be further processed and rounded up or down to directly correspond to a specific OCT B-scan.
[0072] The trained model formed after step 306 may be, for example, model 120 of FIG. 1 . Further, the trained model may be, for example, the model described with respect to process 200 of FIG. 2 . The trained model may be used to accurately and reliably locate the center of the fovea so that other quantitative measurements may be accurately and reliably calculated based on the foveal center location generated by the trained model. For example, the foveal center location may be used to generate an output, such as output 132 described with respect to FIG. 1 . The output may include various retinal thickness measurements associated with CST measurements and / or a retinal grid (e.g., an ETDRS grid). In one or more embodiments, the foveal center location may be used to improve segmentation of retinal layers or fluid features.
[0073] As described above, the training dataset used to train the model may include various types of OCT volumes. In one or more embodiments, the training dataset may include OCT volumes of retinas that have all been diagnosed with a particular retinal disease (e.g., nAMD or diabetic macular edema). In some embodiments, the training dataset includes OCT volumes that have been divided into training OCT volumes, validation OCT volumes, and test OCT volumes. In some embodiments, the OCT volumes in the training dataset are divided into only training OCT volumes and validation OCT volumes.
[0074] IV. Exemplary Workflow for Processing OCT Volumes FIG. 4 is a diagram of an exemplary workflow for processing an OCT volume, according to one or more exemplary embodiments. The OCT volume 400 may be an example of an implementation of the OCT volume 114 of FIG. 1. The OCT volume 400 has an x-, y-, and z-coordinate system, with the x-axis being the lateral axis, the y-axis being the y-axis, and the z-axis being the transverse axis. Each OCT B-scan that makes up the OCT volume 400 may be along, correspond to, or be indexed with a different coordinate on the transverse axis. The OCT volume 400 may have a grayscale similar to that shown in FIG. 4. In other embodiments, some other type of grayscale or black-and-white scale may be used.
[0075] The OCT volume 400 may be preprocessed to form a modified OCT volume 402. The preprocessing may include, for example, but not limited to, at least one of normalization, scaling, resizing, horizontal flipping, vertical flipping, cropping, rotation, noise filtering, or some other type of preprocessing operation. In some examples, the modified OCT volume 402 is created to ensure that the input sent to the model 404 is consistent or substantially similar in size, scale, and / or orientation to the multiple types of OCT volumes on which the model 404 was trained.
[0076] Model 404 may be an example of an implementation of model 120 of FIG. 1. Model 404 includes a 3D convolutional neural network 406 and a recurrent layer 408. 3D convolutional neural network 406 may be an example of an implementation of 3D convolutional neural network 122 of FIG. 1. Recurrent layer 408 may be an example of an implementation of a layer in set of output layers 124 of FIG. 1. Recurrent layer 408 may be considered part of 3D convolutional neural network 406 or may be separate. 3D convolutional neural network 406 of model 404 may be a fully convolutional neural network including a convolutional layer. In one or more embodiments, model 404 includes one or more convolutional layers, one or more subsampling layers, and one or more fully connected layers.
[0077] In one or more embodiments, the hyperparameters of the model 404 include a batch size of 1 and epochs of 60. The model 404 may be implemented using an AdamW optimizer. In one or more embodiments, the model 404 may be specifically designed such that the total number of parameters used in the model 404 is less than approximately 500,000 parameters. The model 404 may be trained using a loss function. The loss function may be, for example, but not limited to, mean absolute error (MAE), mean squared error (MSE), or some other type of loss function. Various metrics may be calculated during training and / or regular use of the model 404. Such metrics include flat R 2 , uniform R 2 , variance-weighted R 2 , Raw R 2 , and / or other types of metrics. In some embodiments, the model 404 may be trained such that the output of the 3D convolutional neural network 406 is normalized (e.g., pixel values are normalized to values between 0 and 1, etc.).
[0078] Model 404 processes corrected OCT volume 402 to generate foveal center location 410. Foveal center location 410 includes three coordinates: the x-coordinate, the y-coordinate, and the z-coordinate of the fovea. In another example, foveal center location 410 includes two coordinates—the x-coordinate and the z-coordinate—for the center of the fovea. Foveal center location 410 may be an example of an implementation of foveal center location 126 described with respect to FIG. 1 and / or an example of an implementation of a foveal center location described with respect to process 200 of FIG. 2, process 300 of FIG. 3, or both.
[0079] In one or more embodiments, the foveal center location 410 is used to generate a CST measurement 412. The CST measurement may be one example of an implementation for the CST measurement 134 described with respect to FIG.
[0080] V. Exemplary Computing System 5 is a block diagram illustrating an example of a computing system 500, according to one or more exemplary embodiments. The computing system 500 may be used to implement the computing platform 102 of FIG. 1 and / or any of the components therein.
[0081] As shown in FIG. 5, computing system 500 may include processor 510, memory 520, storage device 530, and input / output device 540. Computing system 500 may be one exemplary implementation of analysis system 101 of FIG. 1. Processor 510, memory 520, storage device 530, and input / output device 540 may be interconnected via system bus 550. Processor 510 is capable of processing instructions for execution within computing system 500. Such executed instructions may implement one or more components of FIG. 5, such as analysis engine 520, client device 530, etc. In some exemplary embodiments, processor 510 may be a single-threaded processor. Alternatively, processor 510 may be a multi-threaded processor. Processor 510 is capable of processing instructions stored in memory 520 and / or storage device 530 to display graphical information for a user interface, such as display system 106 of FIG. 1.
[0082] Memory 520 is a computer-readable medium, such as a volatile or non-volatile medium, that stores information within computing system 500. Memory 520 may store, for example, data structures representing a configuration object database. Storage device 530 may provide persistent storage for computing system 500. Storage device 530 may be a floppy disk drive, a hard disk drive, an optical disk drive, or a tape drive, or other suitable persistent storage means. Storage device 530 may be an exemplary implementation of data storage 104 of FIG. 1. Input / output device 540 provides input and output operations to computing system 500. In some exemplary embodiments, input / output device 540 includes a keyboard and / or a pointing device. In various embodiments, input / output device 540 includes a display device for displaying a graphical user interface.
[0083] According to some demonstrative embodiments, input / output devices 540 may provide input / output operations for network devices. For example, input / output devices 540 may include an Ethernet port or other networking port for communicating with one or more wired and / or wireless networks (e.g., a local area network (LAN), a wide area network (WAN), the Internet).
[0084] In some exemplary embodiments, computing system 500 can be used to execute various interactive computer software applications that can be used for organizing, analyzing, and / or storing various types of data. Alternatively, computing system 500 can be used to execute any type of software application. These applications can be used to perform various functions, such as planning functions (e.g., creating, managing, editing spreadsheet documents, word processing documents, and / or other objects), computing functions, communication functions, etc. Applications can include various add-in functions or can be standalone computing products and / or features. When active within an application, functionality can be used to generate a user interface that is provided via input / output devices 540. The user interface can be generated by computing system 500 and presented to a user (e.g., on a computer screen monitor, etc.).
[0085] One or more aspects or features of the subject matter described herein may be implemented in digital electronic circuitry, integrated circuits, specially designed ASICs, field programmable gate array (FPGA) computer hardware, firmware, software, and / or combinations thereof. These various aspects or features may include implementation in one or more computer programs executable and / or interpretable on a programmable system including at least one programmable processor, which may be special-purpose or general-purpose, coupled to receive data and instructions from, and transmit data and instructions to, a storage system, at least one input device, and at least one output device. The programmable system or computing system may include clients and servers. Clients and servers are generally remote from each other and typically interact through a communications network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
[0086] These computer programs, sometimes referred to as programs, software, software applications, applications, components, or code, contain machine instructions for a programmable processor and may be implemented in a high-level procedural and / or object-oriented programming language and / or in an assembly / machine language. As used herein, the term “machine-readable medium” refers to any computer program product, apparatus, and / or device used to provide machine instructions and / or data to a programmable processor, such as, for example, magnetic disks, optical disks, memory, and programmable logic devices (PLDs), including a machine-readable medium that receives machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor. A machine-readable medium may non-transitory store such machine instructions, such as, for example, a non-transitory solid-state memory, a magnetic hard drive, or any equivalent storage medium. Alternatively or additionally, a machine-readable medium may temporarily store such machine instructions, such as, for example, a processor cache or other random access memory associated with one or more physical processor cores.
[0087] To provide for user interaction, one or more aspects or features of the subject matter described herein may be implemented on a computer having, for example, a display device, such as a cathode ray tube (CRT) or liquid crystal display (LCD) or light-emitting diode (LED) monitor, for displaying information to a user, and a keyboard and pointing device, such as a mouse or trackball, by which a user may provide input to the computer. Other types of devices may also be used to provide for user interaction. For example, feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, tactile feedback, etc., and input from the user may be received in any form, including acoustic input, voice input, and tactile input. Other possible input devices include touchscreens or other touch-sensitive devices, such as single-point or multi-point resistive or capacitive trackpads, voice recognition hardware and software, optical scanners, optical pointers, digital image capture devices and associated interpretation software, etc.
[0088] VI. Definitions and Contextual Examples The present disclosure is not limited to these exemplary embodiments and applications or the manner in which the exemplary embodiments and applications operate or are described herein. Further, the figures may show simplified or partial views, and the dimensions of the elements in the figures may be exaggerated or out of proportion.
[0089] Unless otherwise defined, scientific and technical terms used in connection with the present teachings described herein shall have the meanings commonly understood by those of ordinary skill in the art. Furthermore, unless the context requires otherwise, singular terms shall include the plural and plural terms shall include the singular. Generally, the nomenclature utilized in connection with, and techniques of, chemistry, biochemistry, molecular biology, pharmacology, and toxicology described herein are those well known and commonly used in the art.
[0090] As used herein, "substantially" means sufficient to function for its intended purpose. Thus, the term "substantially" allows for slight, insignificant variations from an absolute or perfect state, dimension, measurement, result, etc., as would be expected by one of ordinary skill in the art, but does not noticeably affect overall performance. With respect to a parameter or characteristic that is a number or can be expressed as a number, "substantially" means within 10%.
[0091] As used herein, the term "about" when used in reference to a numerical value or a parameter or characteristic that can be expressed as a numerical value means within 10% of the numerical value. For example, "about 50" means a value in the range of 45 to 55.
[0092] The term "ones" means two or more.
[0093] As used herein, the term "plurality" can be 2, 3, 4, 5, 6, 7, 8, 9, 10 or more.
[0094] As used herein, the term "set" means one or more. For example, a set of items includes one or more items.
[0095] As used herein, the phrase "at least one of," when used in conjunction with a list of items, means that different combinations of one or more of the listed items may be used, or that only one of the items in the list may be required. An item may be a specific object, thing, step, action, process, or category. In other words, "at least one of" means that any combination or number of items from the list may be used, but not all of the items in the list may be required. For example, without limitation, "at least one of item A, item B, or item C" means item A; item A and item B; item B; item A, item B, and item C; item B and item C; or items A and C. In some cases, "at least one of item A, item B, or item C" means, but is not limited to, two of item A, one of item B, and ten of item C; four of item B and seven of item C; or some other suitable combination.
[0096] When a reference is made to a list of elements (e.g., elements a, b, c), such reference is intended to include any one of the listed elements alone, any combination of fewer than all of the listed elements, and / or all combinations of the listed elements.
[0097] The terms "one or more of A and B" and "and / or" may also appear in lists of two or more elements or features. Unless otherwise stated, implicitly or explicitly, by the context of use, such phrases are intended to refer to any of the listed elements or features individually, or any of the listed elements or features in combination with any of the other listed elements or features. For example, the phrases "one or more of A and B" and "A and / or B" are intended to mean "A only, B only, or A and B together," respectively. A similar interpretation is intended for lists containing more than two items. For example, the phrases "one or more of A, B, C" and "A, B, and / or C" are intended to mean "A only, B only, C only, A and B together, A and C together, B and C together, or A, B, and C together," respectively.
[0098] The term "based on" means "based at least in part on," meaning that unrecited features or elements are permitted.
[0099] As used herein, a "model" includes at least one of an algorithm, a formula, a mathematical technique, a machine algorithm, a probability distribution or model, a model layer, a machine learning algorithm, or another type of mathematical or statistical representation.
[0100] The term "subject" can refer to a subject of a clinical trial, a person or animal undergoing treatment, a person or animal receiving anti-cancer therapy, a person or animal being monitored for remission or recovery, a person or animal undergoing preventative health analysis (e.g., due to their medical history), or any other person or patient or animal of interest. In various instances, "subject" and "patient" may be used interchangeably herein.
[0101] The term "OCT image" may refer to an image of a tissue, organ, etc., such as the retina, scanned or acquired using optical coherence tomography (OCT) imaging technology. The term may refer to one or both of a 2D "slice" image and a 3D "volume" image. Unless explicitly stated, the term may be understood to include an OCT volume image.
[0102] As used herein, "machine learning" can include the practice of using algorithms to analyze data, learn from it, and then make decisions or predictions about something in the world. Machine learning uses algorithms that can learn from data without relying on rule-based programming.
[0103] As used herein, "artificial neural network" or "neural network" may refer to a mathematical algorithm or computational model that mimics an interconnected group of artificial neurons that process information based on a connectionist approach to computation. A neural network, sometimes referred to as a neural net, may use one or more layers of nonlinear units to predict an output for a received input. Some neural networks include one or more hidden layers in addition to an output layer. The output of each hidden layer is used as the input to the next layer in the network, i.e., the next hidden layer or the output layer. Each layer of the network generates an output from the received input according to the current value of each parameter set. In various embodiments, a reference to a "neural network" may be a reference to one or more neural networks.
[0104] Neural networks, for example, may process information in two ways: they are in training mode when they are being trained (e.g., using a training data set), and they are in inference (or prediction) mode when they are practicing what they have learned (e.g., using a test data set). Neural networks may learn through a feedback process (e.g., backpropagation) that allows the network to adjust the weight coefficients of individual nodes in intermediate hidden layers (modify the behavior of individual nodes) so that their outputs match those of the training data. In other words, neural networks may learn by being fed training data (training examples), and eventually learn how to arrive at the correct output, even when presented with a new range or set of inputs.
[0105] VII. Enumeration of Exemplary Embodiments Embodiment 1. A method comprising: receiving an optical coherence tomography (OCT) volume of a subject's retina, the OCT volume comprising multiple OCT B-scans of the retina; generating a three-dimensional image input of a model using the OCT volume, the model comprising a three-dimensional convolutional neural network; and generating, via the model, a foveal center location comprising three-dimensional coordinates of the center of the fovea of the retina based on the three-dimensional image input.
[0106] Embodiment 2. The method of embodiment 1, wherein the three-dimensional coordinate of the foveal center location includes a transverse coordinate, an axial coordinate, and a lateral coordinate relative to a selected coordinate system of the OCT volume.
[0107] Embodiment 3. The method of embodiment 2 or 3, further comprising using the three-dimensional coordinates of the foveal center location to generate a measurement of the thickness of the central subfield.
[0108] Embodiment 4. The method of any one of embodiments 1 to 3, further comprising determining a retinal grid that divides a retina into regions based on a three-dimensional coordinate of a foveal center position.
[0109] Embodiment 5. The method of embodiment 4, wherein the retinal grid is an Early Treatment Diabetic Retinopathy Study (ETDRS) grid that divides the retina into nine regions centered about the three-dimensional coordinate of the foveal center location.
[0110] Embodiment 6. The method of any one of embodiments 1 to 5, further comprising modifying the segmentation of at least one OCT B-scan of the plurality of OCT B-scans of the retina based on the three-dimensional coordinates of the foveal center location.
[0111] Embodiment 7. The method of any one of embodiments 1 to 6, wherein the model further comprises a recurrent layer used to transform the output of the 3D convolutional neural network into 3D coordinates.
[0112] Embodiment 8. A method according to any one of embodiments 1 to 7, wherein generating a three-dimensional image input includes performing a set of preprocessing operations on the OCT volume to form the three-dimensional image input, the set of preprocessing operations including at least one of a normalization operation, a scaling operation, a resizing operation, a horizontal flip operation, a vertical flip operation, a cropping operation, a rotation operation, or a noise filtering operation.
[0113] Embodiment 9. The method of any one of embodiments 1 to 8, further comprising transforming the three-dimensional coordinates of the foveal center position from a first selected coordinate system associated with the OCT volume to a second selected coordinate system associated with the retina or the subject.
[0114] Embodiment 10 The method of any one of embodiments 1 to 9, wherein the subject's retina is a healthy retina.
[0115] Embodiment 11. The method of any one of embodiments 1 to 9, wherein the subject's retina is diagnosed with age-related macular degeneration (AMD), neovascular age-related macular degeneration (nAMD), diabetic retinopathy, macular edema, or geographic atrophy.
[0116] Embodiment 12. The method of any one of embodiments 1 to 11, wherein one of the three-dimensional coordinates of the foveal center position corresponds to a particular B-scan among multiple OCT B-scans of the retina.
[0117] Embodiment 13. The method of any one of embodiments 1 to 12, further comprising rounding the horizontal coordinate of the three-dimensional coordinate of the foveal center location to a value corresponding to an index associated with a particular OCT B-scan of multiple OCT B-scans of the retina.
[0118] Embodiment 14. The method of any one of embodiments 1 to 12, wherein the foveal center location is generated via a model including a three-dimensional convolutional neural network, and the foveal center location includes rounding an initial value of a horizontal coordinate of the foveal center location to a rounded value corresponding to an index associated with a particular OCT B-scan among a plurality of OCT B-scans of the retina, the rounded value being one of the three-dimensional coordinates of the foveal center location.
[0119] Embodiment 15. A method for training a model, the method comprising: receiving a training dataset including a plurality of optical coherence tomography (OCT) volumes for a plurality of retinas, each of the plurality of OCT volumes including a plurality of OCT B-scans; generating a three-dimensional training image input for a model using the plurality of OCT volumes of the training dataset, the model including a three-dimensional convolutional neural network and a recurrent layer; and training the model to generate a foveal center location including three-dimensional coordinates of the center of the fovea of the retina of a selected OCT volume based on the three-dimensional training image input.
[0120] Embodiment 16. The method of embodiment 15, wherein the foveal center location comprises the transverse, lateral, and axial coordinates of the center of the fovea in the OCT volume.
[0121] Embodiment 17. The method of embodiment 15 or embodiment 16, wherein generating a training 3D image input includes performing a set of preprocessing operations on a plurality of OCT volumes to form the training 3D image input, and the set of preprocessing operations includes at least one of a normalization operation, a scaling operation, a resizing operation, a horizontal flip operation, a vertical flip operation, a cropping operation, a rotation operation, or a noise filtering operation.
[0122] Embodiment 18. The method of any one of embodiments 15 to 17, wherein the plurality of retinas comprises at least one healthy retina.
[0123] Embodiment 19. The method of any one of embodiments 15 to 18, wherein the plurality of retinas comprises at least one retina diagnosed with a retinal disease that is age-related macular degeneration (AMD), neovascular age-related macular degeneration (nAMD), diabetic retinopathy, macular edema, or geographic atrophy.
[0124] Embodiment 20. The method of any one of embodiments 15 to 19, wherein one of the three-dimensional coordinates of the foveal center position corresponds to a particular B-scan among multiple OCT B-scans of the retina.
[0125] Embodiment 21. A system including one or more data processors and a non-transitory computer-readable storage medium including instructions, which, when executed on the one or more data processors, cause the one or more data processors to perform the following: receive an optical coherence tomography (OCT) volume of a subject's retina, the OCT volume including a plurality of OCT B-scans of the retina; generate a three-dimensional image input of a model using the OCT volume, the model including a three-dimensional convolutional neural network; and generate, via the model, a foveal center location including three-dimensional coordinates of the center of the fovea of the retina based on the three-dimensional image input.
[0126] Embodiment 22. A computer program product tangibly embodied in a non-transitory machine-readable storage medium containing instructions, the instructions configured to cause one or more data processors to perform the following: receive an optical coherence tomography (OCT) volume of a subject's retina, the OCT volume including a plurality of OCT B-scans of the retina; generate a three-dimensional image input for a model using the OCT volume, the model including a three-dimensional convolutional neural network; and generate, via the model, a foveal center location including three-dimensional coordinates of the center of the fovea of the retina based on the three-dimensional image input.
[0127] Embodiment 23. A computer program product tangibly embodied in a non-transitory machine-readable storage medium comprising instructions configured to cause one or more data processors to perform the method of any one or more of embodiments 1 to 20.
[0128] Embodiment 24. A system including one or more data processors and a non-transitory computer-readable storage medium containing instructions, which, when executed on the one or more data processors, cause the one or more data processors to perform a method according to any one or more of embodiments 1 to 20.
[0129] VIII. Further Considerations While the present teachings will be described in conjunction with various embodiments, it is not intended that the present teachings be limited to such embodiments. On the contrary, the present teachings encompass various alternatives, modifications, and equivalents, as will be appreciated by those skilled in the art.
[0130] In describing various embodiments, the specification may present a method and / or process as a particular sequence of steps. However, to the extent that the method or process does not rely on the particular order of steps set forth herein, the method or process should not be limited to the particular order of steps set forth, and as one skilled in the art will readily appreciate, the order may be changed and still remain within the spirit and scope of the various embodiments.
[0131] Furthermore, the subject matter described herein may be embodied in systems, devices, methods, and / or articles, depending on the desired configuration. The implementations set forth in the description herein do not necessarily represent all implementations of the described subject matter. Rather, they are merely some examples consistent with aspects related to the described subject matter. While several variations have been detailed above, other modifications and additions are possible. In particular, additional features and / or variations may be provided in addition to those described herein. For example, the implementations described above may be directed to various combinations and subcombinations of the disclosed features and / or combinations and subcombinations of several additional features disclosed above. Additionally, the logic flow depicted in the accompanying figures and / or described herein does not necessarily require the particular order shown or sequential order to achieve desirable results. Other implementations may be within the scope of the following claims.
Claims
1. 1. A method comprising: receiving an optical coherence tomography (OCT) volume of a subject's retina, the OCT volume including a plurality of OCT B-scans of the retina; generating a three-dimensional image input of a model using the OCT volume, the model including a three-dimensional convolutional neural network; and generating, via the model, a foveal center position including three-dimensional coordinates of the center of the fovea of the retina based on the three-dimensional image input; A method comprising:
2. The method of claim 1 , wherein the three-dimensional coordinates of the foveal center location include a transverse coordinate, an axial coordinate, and a lateral coordinate relative to a selected coordinate system of the OCT volume.
3. The method of claim 2 or 3, further comprising using the three-dimensional coordinates of the foveal center location to generate a central subfield thickness measurement.
4. The method of claim 1 , further comprising determining a retinal grid that divides the retina into regions based on the three-dimensional coordinates of the foveal center location.
5. 5. The method of claim 4, wherein the retinal grid is an Early Treatment Diabetic Retinopathy Study (ETDRS) grid that divides the retina into nine regions centered about the three-dimensional coordinate of the foveal center location.
6. 6. The method of claim 1, further comprising modifying a segmentation of at least one OCT B-scan of the plurality of OCT B-scans of the retina based on the three-dimensional coordinates of the foveal center location.
7. 7. The method of claim 1, wherein the model further comprises a recurrent layer used to transform the output of the 3D convolutional neural network into the 3D coordinates.
8. generating the three-dimensional image input comprises:
8. The method of claim 1, further comprising performing a set of pre-processing operations on the OCT volume to form the three-dimensional image input, the set of pre-processing operations comprising at least one of a normalization operation, a scaling operation, a resizing operation, a horizontal flip operation, a vertical flip operation, a cropping operation, a rotation operation, or a noise filtering operation.
9. 9. The method of claim 1, further comprising transforming the three-dimensional coordinates of the foveal center location from a first selected coordinate system associated with the OCT volume to a second selected coordinate system associated with the retina or the subject.
10. 10. The method of claim 1, wherein the retina of the subject is a healthy retina.
11. 10. The method of any one of claims 1 to 9, wherein the subject's retina is diagnosed with age-related macular degeneration (AMD), neovascular age-related macular degeneration (nAMD), diabetic retinopathy, macular edema, or geographic atrophy.
12. The method of claim 1 , wherein one of the three-dimensional coordinates of the foveal center position corresponds to a particular B-scan of the plurality of OCT B-scans of the retina.
13. 13. The method of claim 1, further comprising: rounding a horizontal coordinate of the three-dimensional coordinates of the foveal center location to a value corresponding to an index associated with a particular OCT B-scan of the plurality of OCT B-scans of the retina.
14. generating the foveal center location via the model including the 3D convolutional neural network; 13. The method of claim 1, further comprising: rounding an initial value of a horizontal coordinate of the foveal center location to a rounded value corresponding to an index associated with a particular OCT B-scan of the plurality of OCT B-scans of the retina, the rounded value being one of the three-dimensional coordinates of the foveal center location.
15. 1. A method for training a model, comprising: receiving a training dataset comprising a plurality of optical coherence tomography (OCT) volumes for a plurality of retinas, each of the plurality of OCT volumes comprising a plurality of OCT B-scans; generating a training 3D image input for a model using the plurality of OCT volumes of the training dataset, the model including a 3D convolutional neural network and a recurrent layer; and training the model to generate a foveal center location including three-dimensional coordinates of the center of the fovea of the retina of a selected OCT volume based on the training three-dimensional image input; A method comprising:
16. The method of claim 15 , wherein the foveal center location comprises the transverse, lateral, and axial coordinates of the center of the fovea of the OCT volume.
17. generating the training three-dimensional image input comprises:
17. The method of claim 15 or claim 16, comprising performing a set of pre-processing operations on the plurality of OCT volumes to form the training three-dimensional image input, the set of pre-processing operations comprising at least one of a normalization operation, a scaling operation, a resizing operation, a horizontal flip operation, a vertical flip operation, a cropping operation, a rotation operation, or a noise filtering operation.
18. 18. The method of any one of claims 15 to 17, wherein the plurality of retinas includes at least one healthy retina.
19. 19. The method of any one of claims 15 to 18, wherein the plurality of retinas includes at least one retina diagnosed with a retinal disease that is age-related macular degeneration (AMD), neovascular age-related macular degeneration (nAMD), diabetic retinopathy, macular edema, or geographic atrophy.
20. 20. The method of claim 15, wherein one of the three-dimensional coordinates of the foveal center location corresponds to a particular B-scan of the plurality of OCT B-scans of the retina.
21. 1. A system comprising: one or more data processors; 1. A non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to: receiving an optical coherence tomography (OCT) volume of a subject's retina, the OCT volume including a plurality of OCT B-scans of the retina; generating a three-dimensional image input of a model using the OCT volume, the model including a three-dimensional convolutional neural network; and generating, via the model, a foveal center position including three-dimensional coordinates of the center of the fovea of the retina based on the three-dimensional image input; a non-transitory computer-readable storage medium for executing the Including, the system.
22. A computer program product tangibly embodied in a non-transitory machine-readable storage medium containing instructions, said instructions causing one or more data processors to: receiving an optical coherence tomography (OCT) volume of a subject's retina, the OCT volume including a plurality of OCT B-scans of the retina; generating a three-dimensional image input of a model using the OCT volume, the model including a three-dimensional convolutional neural network; and generating, via the model, a foveal center position including three-dimensional coordinates of the center of the fovea of the retina based on the three-dimensional image input; A computer program product configured to cause the computer to execute