OCT-EN-FACE Lesion Segmentation Using Channel-Coded Slabs

JP7791106B2Active Publication Date: 2025-12-23CARL ZEISS MEDITEC INC +1
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2022564330
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-04-29
Filing Date
2021-04-28
Publication Date
2025-12-23
Estimated Expiration
2041-04-28

Smart Images

  • Figure 0007791106000001
    Figure 0007791106000001
  • Figure 0007791106000002
    Figure 0007791106000002
  • Figure 0007791106000003
    Figure 0007791106000003
Patent Text Reader

Abstract

A system and method for identifying target lesions using optical coherence tomography (OCT) data extracts multiple lesion characteristic images from the OCT data. The extracted lesion characteristic images may include a blend of an OCT structural image (including retinal layer thickness information) and an OCT angiography image. Optionally, other lesion characteristic images and data maps (mapped to corresponding locations in the OCT data), such as fundus images and visual field test maps, may be accessed as additional lesion characteristic images. Each lesion characteristic image defines a different image channel (e.g., a "color channel") for each pixel in the composite channel-encoded image, which is then used to train a neural network to search for target lesions in the OCT data. The trained neural network may then receive a new composite channel-encoded image and identify / segment the target lesion in the new channel-encoded image.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates generally to methods of analyzing optical coherence tomography data to identify target pathology, and more particularly to analyzing optical coherence tomography data to identify geographic atrophy (GA). [Background technology]

[0002] Age-related macular degeneration (AMD) is the most common eye disease in older adults. It results from damage to the macula, leading to loss of central vision. Some patients with AMD develop geographic atrophy (GA), which refers to an area of ​​the retina where cells thin and die. GA is a macular disorder that appears in the advanced stages of nonexudative macular degeneration. GA has a distinctive appearance resulting from defects in the photoreceptor layer, retinal pigment epithelium (RPE), and choriocapillaris. GA typically first appears in a parafoveal location, progresses around the fovea, and then progresses through the fovea, accompanied by loss of central vision. While there are currently no known treatments that effectively delay or reverse the effects of GA, characterizing and monitoring the macular area affected by GA is fundamental for patient diagnosis, monitoring, and management, as well as for therapeutic research purposes.

[0003] The appearance of GA has been studied in reflectance (color) fundus imaging, autofluorescence imaging, and more recently, optical coherence tomography (OCT) imaging (e.g., using 2D topographic imaging techniques). OCT effectively visualizes areas of GA not by directly visualizing RPE disruption but by exploiting the effect of this disruption on light transmitted through the choroid. The RPE is a highly reflective layer for OCT signal, and the increased transmission of light (OCT signal) to the atrophic choroid allows visualization of the presence of GA in en face sub-RPE (sub-RPE) images, which can be formed by axial projection of a subvolume (e.g., a slab) of OCT data extending from the RPE (or slightly above the RPE) to below the RPE and into the choroid. The presence of GA in OCT images can be identified as brighter areas in the en face sub-RPE images. This characteristic of GA in OCT images can be referred to as sub-RPE hyperreflectivity.

[0004] Quantification of GA characteristics (such as area size or distance to the foveal center, as can be seen in en face RPE-underlying images or other 2D en face images) that may be valuable or interesting for monitoring this condition relies on delineating or segmenting the GA area within the macula. However, manually segmenting the GA area in OCT data is a difficult and time-consuming task.

[0005] While sub-RPE hyperreflectivity remains a valid technique for visualizing GA in OCT data, the presence of GA is inferred (e.g., based on variations in reflectivity) but not directly observed. As a result, this technique is subject to potential difficulties arising from other factors that can also affect reflectivity (e.g., the presence of superficial or choroidal vessels that cast shadows, possible retinal opacities as hyperreflective lesions, or areas of increased choroidal signal within an intact RPE). Due to these difficulties, en face (e.g., frontal, planar view) images of the sub-RPE region are not sufficient, and careful B-scan review of the suspected GA area is required to confirm its presence. B-scans provide axial slice (or lateral) views of the region (e.g., slices through the retinal layers), and the presence of GA can be confirmed by noting verifiable defects in the RPE layer and retinal thinning or collapse.

[0006] The most common automated GA segmentation methods are based on the analysis of sub-RPE hyperreflectivity in a single sub-RPE en face image and are subject to the difficulties and potential errors outlined above. Because a single sub-RPE en face image does not provide all the necessary information, careful B-scan review was required to check and correct for possible errors. Summary of the Invention [Problem to be solved by the invention]

[0007] What is needed is a method for identifying the presence of GA that takes into account information from the entire OCT volume, but is as simple to use as traditional methods using en face images of the RPE.

[0008] It is an object of the present invention to provide a more accurate GA detection system for use with OCT. Another object of the present invention is to provide an OCT-based GA detection method that takes into account multiple types of OCT information, such as may be obtained from a combination of en face and B-scan images.

[0009] It is a further object of the present invention to provide an OCT-based system that provides an en face image representation of a combination of OCT-based data, non-OCT image data, and / or non-image data.

[0010] It is yet another object of the present invention to provide a system / method for identifying / segmenting GA regions in en face OCT images, taking into account image and non-image information from different types of imaging modalities. [Means for solving the problem]

[0011] The above objectives are achieved in a method / system for analyzing optical coherence tomography (OCT) data to identify / segment target lesions (e.g., geographic atrophy) within en face images. The present invention integrates different aspects of OCT volumetric information into different image channels (e.g., color channels) of a single image to create an improved GA detection, analysis, and segmentation tool. The present method provides better GA visualization in a single image that can be used in segmentation algorithms to provide more accurate segmentation (e.g., GA segmentation), or provides for review and editing of segmentation results. Unlike prior art techniques that focus on displaying and characterizing a single factor associated with the possible presence of GA (e.g., sub-RPE reflection), the present approach integrates multiple different aspects of the OCT volume into additional channels (e.g., pixel / voxel channels) that are combined to more accurately analyze and segment GA. For example, the additional information may include analysis of or data related to RPE integrity or retinal thinning (e.g., thinning of specific retinal layers), which previously required analysis of individual B-scans.

[0012] Other objects and achievements of the present invention, together with a fuller understanding of the invention, will become apparent and understood by reference to the following description and claims taken in conjunction with the accompanying drawings.

[0013] To facilitate the understanding of the present invention, several publications are cited or referenced herein. All publications cited or referenced herein are incorporated by reference in their entirety.

[0014] The embodiments disclosed herein are merely examples, and the scope of the present disclosure is not limited thereto. Features of any embodiment described in one claim category, e.g., a system, may also be claimed in other claim categories, e.g., a method. Dependencies or back-references in the appended claims are selected for formality reasons only. However, any subject matter available from a careful back-reference to a previous claim may also be claimed, thereby disclosing any combination of claims and their features and may be claimed regardless of the dependencies selected in the appended claims. [Brief explanation of the drawings]

[0015] In the drawings, like reference numbers / letters refer to like elements. [Figure 1] FIG. 1 shows three examples of GA in images of three different imaging modalities. [Figure 2] FIG. 10 provides an example of a B-scan through the GA region. [Figure 3] FIG. 1 illustrates an embodiment of the present invention. [Figure 4A] FIG. 1 shows a sub-RPE slab projection, as may conventionally be used to identify candidate GA regions. [Figure 4B] FIG. 1 illustrates a channel-encoded image (multi-channel composite image) according to the present invention, where each image channel (eg, color channel) embodies different lesion-specific characteristics. [Figure 5]FIG. 10 illustrates an alternative embodiment in which a multi-channel composite image is composed of images of different imaging modalities and optionally non-image data. [Figure 6] FIG. 1 illustrates a schematic workflow of the present invention, including the U-net architecture of a proof-of-concept embodiment. [Figure 7] FIG. 10 shows qualitative results of this embodiment of the proof-of-concept. [Figure 8] FIG. 10 shows quantitative measurements of this embodiment of the proof-of-concept. [Figure 9] FIG. 1 provides an overview of the present invention. [Figure 10] FIG. 1 illustrates an exemplary visual field testing device (perimeter) for testing a subject's visual field. [Figure 11] FIG. 1 illustrates an example of a slit-scan ophthalmic system for imaging the fundus. [Figure 12] FIG. 1 illustrates a generalized frequency-domain optical coherence tomography system used to collect 3D image data of the eye suitable for use in the present invention. [Figure 13] FIG. 1 shows an exemplary OCT B-scan image of a normal retina of a human eye, illustratively identifying various normal retinal layers and boundaries. [Figure 14] FIG. 10 shows an example of an en face vascular image. [Figure 15] FIG. 1 shows an exemplary B-scan of a vasculature (OCTA) image. [Figure 16] FIG. 1 illustrates an example of a multi-layer perceptron (MLP) neural network. [Figure 17] FIG. 1 illustrates a simplified neural network consisting of an input layer, a hidden layer, and an output layer. [Figure 18] FIG. 1 illustrates an exemplary convolutional neural network architecture. [Figure 19] FIG. 1 illustrates an exemplary U-Net architecture. [Figure 20] FIG. 1 illustrates an exemplary computer system (or computing device or computer). DETAILED DESCRIPTION OF THE INVENTION

[0016] Geographic atrophy (GA) refers to areas of the retina where cells become thin and die (atrophy). These atrophic areas typically result in blind spots in a person's visual field. Therefore, monitoring and characterization of retinal areas affected by GA is essential for patient diagnosis and management. Various imaging modalities, such as fundus imaging (including autofluorescence and fluorescein angiography), optical coherence tomography (OCT), and OCT angiography (OCTA), have proven useful for detecting and characterizing GA.

[0017] Fundus imaging, such as may be obtained by use of a fundus camera, generally provides a frontal, planar view of the fundus as seen through the pupil of the eye. Fundus imaging may use different frequencies of light, such as white light, red light, blue light, green light, infrared light, etc., to image tissue, or may use selected frequencies to excite fluorescent molecules within specific tissues (e.g., autofluorescence) or to excite fluorescent dyes injected into the patient (e.g., fluorescence angiography). A more detailed description of different fundus imaging techniques is provided below.

[0018] OCT is a noninvasive imaging technique that uses light waves to generate cross-sectional images of retinal tissue. For example, OCT allows for the visualization of distinct tissue layers of the retina. Generally, OCT systems are interferometric imaging systems that determine the scattering profile of a sample along the OCT beam by detecting the interference of light reflected from the sample with a reference beam, forming a three-dimensional (3D) representation of the sample. Each scattering profile in the depth direction (e.g., z-axis or axial direction) can be individually reconstructed into an axial scan or A-scan. Cross-sectional two-dimensional (2D) images (B-scans) and extended 3D volumes (C-scans or cube scans) can be constructed from multiple A-scans acquired as the OCT beam is scanned / translated through a set of transverse (e.g., x-axis and y-axis) locations on the sample. OCT also allows for the construction of planar, en face (e.g., en face) 2D images of selected portions of a tissue volume (e.g., a target tissue slab (subvolume) or target tissue layer(s) of the retina). OCTA is an extension of OCT and can identify (e.g., render in an image format) the presence or absence of blood flow in tissue layers. OCTA can identify blood flow by identifying differences (e.g., contrast differences) over time in multiple OCT images of the same retinal region and designating differences that meet predetermined criteria as blood flow. A more detailed description of OCT and OCTA is provided below.

[0019] Each imaging modality can characterize GA differently. Figure 1 shows three examples of GA in images from three different imaging modalities. Image 11 is a reflectance (color) fundus image, and image 13 is an autofluorescence image (obtained using 2D topographic imaging techniques). Image 15 is an en face, sub-RPE, OCT image. In images 11, 13, and 15, GA is identified as region 17. GA typically results in loss or thinning of some retinal layers (e.g., the photoreceptor layer, the retinal pigment epithelium (RPE), and the choriocapillaris). Suspicious regions in the en face view can be identified, at least in part, by the loss / thinning of these retinal layers caused by GA. The thinning of these layers can result in characteristic color, intensity, or texture changes in the en face, planar view of the retina, as seen in fundus, OCT, and / or OCTA images. For example, in a healthy retina, these layers, particularly the RPE, tend to reflect OCT signals and limit their penetration into the retina. However, in areas of GA, the OCT signal tends to penetrate deeper into the retina due to thinning or loss of these layers, resulting in the characteristic sub-RPE hyperreflectivity, as shown in Image 15. This hyperreflectivity can help identify candidate / suspicious areas that may be GA, but it is not a direct detection of GA, as other factors may also affect OCT signal reflectivity.

[0020] A more direct method of observing GA is to use B-scans, which provide a slice view of the GA region and show individual retinal layers. Figure 2 provides an example of a B-scan that penetrates the GA region. As shown, the GA region 19 can be identified by thinning of the RPE and other layers, increased reflectivity / brightness below the RPE layer, and / or overall thinning or disintegration of the retina. Therefore, a more direct way to search for GA is to examine individual B-scans, but examining each B-scan within a volume for the presence of GA is generally time-consuming. As a result, a more rapid method for detecting GA is to examine a single frontal, planar view of the fundus / retina, identify a suspicious region, such as those shown in images 11, 13, or 15 in Figure 1, and then examine individual B-scans (one by one) that traverse the suspicious region.

[0021] GA detection, characterization, and segmentation in OCT data have traditionally been performed in sub-RPE en face slabs / images by analyzing the increased signal in the choroid derived from RPE disruption. Examples of this technique include Qiang Chen et al., "Semi-automatic geographic atrophy segmentation for SD-OCT images," Biomed. Opt. Express 4, pp. 2729–2750 (2013), and Sijie Niu et al., "Automated geographic atrophy segmentation for SD-OCT images using region-based CV model via local similarity factor," Biomed. Opt. Express 7, pp. 581–600 (2016). However, this technique is subject to potential difficulties and limitations due to any of the many factors that can affect the signal (e.g., reflectance) observed in RPE subface images. As a result, commercially available automated GA segmentation tools based on this technique are known to provide less than optimal results and require subsequent B-scan review to confirm the presence of potential GAs.

[0022] A different approach is described by Ji Z. et al., "Retinal Layers: A Deep Voting Model for Automated Geographic Atrophy Segmentation in SD-OCT Images," Transl Vis Sci Technol., 2018, 7(1). This approach is based on neural networks and directly processes individual B-scans for GA segmentation. However, this process can be complicated by the need to annotate individual B-scans, which can result in discontinuous results in adjacent B-scans. Furthermore, this approach does not solve the problem of requiring subsequent B-scans for each B-scan to confirm the GA segmentation results.

[0023] Therefore, OCT-based GA characterization methods have traditionally focused on analyzing sub-RPE reflectance in en face images, excluding important information such as direct analysis of RPE integrity or retinal thinning. To account for such information, additional review of multiple images (e.g., from different imaging modalities) or multiple OCT B-scans has traditionally been required.

[0024] The present invention collects GA characteristic information from multiple different image views (and / or imaging modalities) and combines or incorporates that information into a single, customized en face image of the retina, thereby providing multiple sources of information to enable more accurate identification of GA regions. That is, the present invention provides several sources of information in a single image that can be used to manually determine the presence of GA or in automated or semi-automated GA segmentation tools. Different characteristic information can be stored (or encoded) in additional channels (e.g., color channels) of the image, such as on a pixel (or voxel) basis. The combination of different characteristics (e.g., GA-related information) into different channels of a single image allows for more accurate and effective characterization and segmentation of GA regions.

[0025] In summary, a first exemplary embodiment of the present invention uses, for example, a set of en face images (front view, topographic projections from an OCT volume forming a 2D image) with different definitions and stacks their information in different channels (e.g., color channels) of a single image to generate a channel-encoded image in which different characteristics characterizing the presence of GA (or other target lesion) are encoded in the different channels. A preliminary step may be segmentation of a set of retinal layers within the OCT volume to aid in en face creation (e.g., to aid in selecting layers that may be characteristic of (e.g., associated with) the target lesion and that may be used to define the slabs from which the en face image is created). A set of different en face images is then created using the segmented layers and / or a set of slab definitions. The en face images may then be stacked (or combined) in different channels (e.g., pixel / voxel color channels). The resulting set of stacked slabs can be used to visualize GA (or other specific pathologies) and other retinal landmarks (e.g., if three separate slabs are stacked instead of the individual red, green, and blue (RGB) color channels of a typical color image), and / or can be used for automatic segmentation of GA using the information encoded in the different channels.

[0026] 3 illustrates one embodiment of the present invention. The first step is acquiring OCT data 21 (e.g., one or more data volumes of the same region of the retina). Optionally, the OCT data 21 may be provided to a retinal layer segmentation process. As part of, or in addition to, this process, each A-scan in the OCT data 21 is processed to extract a set of metrics 23, each metric selected to measure (or emphasize, such as by weighting) a characteristic of a particular target lesion (e.g., GA). That is, the extracted metrics may be associated with the same lesion type, such that each metric provides a different marker for the same lesion type. The extracted metrics are sorted or otherwise collected into corresponding metric groups (metric-1 group through metric-n group), optionally in a one-to-one correspondence. The OCT data 21 can include OCT structural data and / or OCTA flow data, and the extracted metrics can include OCT-based metrics extracted from the OCT structural data, such as retinal layer thickness, distance from a particular A-scan to a particular retinal structure (e.g., distance to the fovea), layer integrity (e.g., defects in a particular layer), sub-RPE reflectivity, internal RPE reflectivity, total retinal thickness, and / or optical attenuation coefficient (OAC).

[0027] The optical attenuation coefficient (OAC) is an optical property of a (e.g., turbid) medium (e.g., tissue) that determines how the power of a coherent light beam propagating through it attenuates along its path due to scattering and absorption. The irradiance (power per unit area) of a coherent light beam propagating through a (e.g., homogeneous) medium is determined by the Lambert-Beer law, L(z) = L0e -μzwhere L(z) is the irradiance of the beam after passing through the medium over distance z, L0 is the irradiance of the incident light beam, and μ is the optical attenuation coefficient. A large attenuation coefficient results in a rapid, exponential drop in the irradiance of a coherent light beam with depth. Because the OAC is an optical property of a medium, determining the OAC provides information about the composition of this medium. Applicant proposes that providing the OAC (per A-scan) as one of the extracted metrics may be useful for identifying certain pathologies (e.g., GA), particularly since it can indicate the current state (e.g., optical attenuation state) of the tissue at a particular A-scan location. An example of how OAC can be determined / calculated is provided in K.A. Vermeer et al., "Depth-Resolved Model-Based Reconstruction of Attenuation Coefficients in Optical Coherence Tomography," Biomedical Optics Express, Vol. 5, No. 1, pp. 322-337 (2014). A description of a previous application of OAC is Utka Baran et al., "In Vivo Tissue Injury Mapping Using Optical Coherence Tomography Based Methods," Applied Optics, Vol. 54, No. 21, July 20, 2015.

[0028] The extracted metrics may also include OCTA-based metrics extracted from the OCTA flow data, such as flow measurements (e.g., blood flow) at locations within one or more layers (e.g., choriocapillaris, Satler's layers, Haller's layer, etc.) and distance from the flow data to the fovea. In this manner, each metric group can describe a different lesion characteristic, and multiple metric groups (metric-1 group through metric-n group) can be used to define multiple corresponding lesion characteristic images (PCI-1 through PCI-n), each highlighting a different lesion characteristic. Each lesion characteristic image PCI-1 through PCI-n may define an en face image. The different lesion characteristic images may then be used to define different pixel channels (Ch1 through Chn) and combined (as indicated by block 25) to form a multi-channel composite image (e.g., a channel-encoded image) 27. Optionally, the multi-channel composite image 27 may be of lower dimensionality than the OCT data 21, such as an en face image (and / or a B-scan image), with each pixel location in the composite image 27 based on a corresponding A-scan location in the OCT data 21. In this manner, each metric may be used as the basis for a different corresponding channel in the composite image 27. The resulting multi-channel composite image 27 may then be provided to a machine learning model 29 trained to identify target lesions (e.g., GAs) based on the lesion characteristic data (e.g., metric groups) embodied in the individual image channels. The identified lesions may then be displayed or stored for future processing in the computing device 31.

[0029] As an example, a proof-of-concept embodiment used three metric groups to define three different slabs (lesion characteristic images) assigned to the three typical red, green, and blue (RGB) color channels of an image. It should be understood that a channel-coded image may optionally have more (or fewer) channels. Figure 4A shows a sub-RPE slab projection as might conventionally be used to identify candidate GA regions, while Figure 4B shows a channel-coded image (multi-channel composite image) according to the present invention, where each image channel (e.g., color channel) embodies a different lesion-specific characteristic (embodies a different metric group). In this example, the red channel (or mid-gray in a black-and-white monochrome image) contains sub-RPE reflectance data. To collect metrics for the red channel, a 300 μm slab is defined outside the RPE layer and near the choroid, with surface limits specified between the RPE layer and plus offsets of 50 μm and 350 μm, respectively. The OCT signal within this slab is filtered to remove noise and then processed so that the signal at each A-scan location has a constantly decreasing function with increasing depth, filling in any "valleys" in the signal. That is, for each particular pixel in the A-scan, a value is set as the highest value recorded in that A-scan from the pixel under consideration to increasing depths within the defined slab limits. This processing is set to remove lower-value signals due to the presence of choroidal vessels. The resulting data is projected into an en face image by averaging the pixel values ​​within the slab limits for each A-scan. The values ​​of the resulting en face image are then normalized to fall within a range between 0 and 1. The goal of this slab is to characterize the increased reflectivity present in the choroid in the GA region.

[0030] In this example, the green channel (e.g., light gray in a black-and-white monochrome image) encompasses RPE internal reflectance. To collect metrics for the green channel, a 20 μm slab is defined inside the RPE-Fit layer (an estimate of Bruch's membrane curvature set at the level of the RPE centerline), and surface limits are specified between the RPE-Fit layer and offsets of minus 50 μm and minus 30 μm, respectively. The OCT signal within this slab is filtered for noise removal and then processed so that the signal at each A-scan location has a constantly increasing function with increasing depth, filling in "valleys" in the signal. That is, for each specific pixel in the A-scan, a value is set as the highest value recorded in such an A-scan from the inner slab limit to the pixel under consideration. This processing is set to remove lower-value signals caused by shadowing of higher-opacity structures (e.g., blood vessels, drusen, or hyperreflective lesions) in the otherwise intact RPE. The resulting data is projected into an en face image by averaging pixel values ​​within the slab limits for each A-scan. The values ​​of the resulting en face image are normalized to fall within a range between 0 and 1. The goal of this slab is to characterize the lower reflectance in areas with photoreceptor and RPE defects.

[0031] In this example, the blue channel (or dark gray in a black-and-white monochrome image) encompasses retinal thickness. To collect metrics for the blue channel, the distance between the ILM layer and the RPE-fit layer (retinal thickness) is measured for each A-scan location and projected onto the en face image. The recorded values ​​are then scaled with an inverse linear operation to take values ​​between 0 and 1, such that a retinal thickness of 100 μm takes a value of 1 and a thickness of 350 μm takes a value of 0. The goal of this slab is to characterize localized areas of retinal thinning and collapse characteristic of the presence of GA.

[0032] FIG. 5 illustrates an alternative embodiment in which a multi-channel composite image (i.e., a color-coded image or a monochrome-coded image) 27 is composed of images of different imaging modalities and, optionally, non-image data. In this embodiment, OCT data 21 is acquired or otherwise accessed. The OCT data 21 may be a cube scan composed of multiple A-scans and / or accessed B-scans, and may optionally include multiple scans of the same region separated in time. In this case, OCTA data 22 may be determined from the OCT data 21, as indicated by the dotted arrow 24. Alternatively, the OCTA data 22 may be acquired / accessed separately. A set of metrics may then be extracted from each of the OCT data 21 and the OCTA data 22, and a set of lesion characteristic images (or maps) may be defined from the extracted metrics (e.g., metric groups). In this example, images OCT1-OCT4 are defined based on the metrics extracted from the OCT data 21, and images OCTA1-OCTA3 are defined based on the metrics extracted from the OCTA data 22. For example, images OCT1-OCT4 may represent, respectively, an en face RPE subimage, a thickness map of a select layer (e.g., the photoreceptor layer, the retinal pigment epithelium (RPE), and / or the choriocapillaris, which may further localize the fovea, such as by use of an automated foveal localization algorithm), a layer integrity map (e.g., as typically determined from multiple B-scans), and an OAC map. In this example, images OCTA1-OCTA3 may represent, respectively, an en face OCTA image of choriocapillaris flow, an image of Sattler's layer (e.g., located between Bruch's membrane, the choriocapillaris, and the Haller layer below and the suprachoroidea above), and / or flow maps of other select layers relative to the foveal location.

[0033] As explained above, GA can result in a progressive loss of vision, particularly central vision. However, GA can begin with a loss of vision outside the central area and progress toward the center over time. Therefore, it is advantageous to incorporate information from visual field test results (FV). Visual field testing is a method of measuring an individual's entire visual field, for example, central and peripheral (side) vision. Visual field testing is a method of individually mapping the visual field of each eye and can detect blind spots (scotoma) and more subtle areas of mesopic vision. A campimeter, or "perimeter," is a dedicated instrument / device / system that administers visual field tests to patients. A more detailed description of perimeters and visual field tests is provided below. All or selected portions of the visual field test (such as VF grayscales or numeric grayscales mapped to corresponding retinal locations) can be incorporated into the present multi-channel composite image 27.

[0034] Additional imaging modalities may include one or more types of fundus images FI (e.g., white light, red light, blue light, green light, infrared, autofluorescence, etc.) and fluorescence angiography image(s) FL.

[0035] Each of the different data types described above may represent a different lesion characteristic image and may be combined, as indicated by block 25, to define a multi-channel composite image 27. As shown, each pixel (shown as a circle Pxl) may contain data (e.g., metrics) from each of the above sources. For example, each pixel may define a data record consisting of multiple data fields (for the corresponding retinal location), one for each incorporated lesion characteristic image. Each pixel may include a visual field test data field (VF-1), a fundus image data field (FI-1), a fluorescein angiography image data field (FL-1), OCT structural data fields (OCT1-1, OCT2-1, OCT3-1, and OCT1-4) from each corresponding OCT structural image, and an OCTA flow data field (OCTA1-1, OCTA2-1, and OCTA3-1) from each corresponding OCTA flow image.

[0036] The composite image 27 may then be provided to a machine learning model 29 for processing or training, as described below. As in the embodiment of FIG. 5, output from the machine learning model 27 may be provided to a computing device (not shown) for display or storage. Optionally, non-image data 28 may also be provided to a machine model 29 for processing, either indirectly via dotted arrow 26A or directly via dotted block arrow 26B. That is, the non-image data 28 may be optimally incorporated into the composite image 27 via block 25. The non-image data 28 may include patient demographic data (e.g., age, ethnicity, etc.) and / or medical history data (e.g., previously prescribed medication(s) and diagnosed illnesses related to the sought-after pathology, etc.), such as may be obtained from an electronic medical record (EMR) system.

[0037] A proof-of-concept application of the present invention implements a machine learning model 29 as a neural network architecture trained for the automatic segmentation of GA region(s) in a composite image 27 (e.g., a generated channel-coded image). A schematic description of the neural network is provided below. All accessed images (and / or maps) used to define the composite image 27 may be normalized and resized to 256 x 256 x 3 pixels. Each image is then divided into nine overlapping patches of pixel size 128 x 128 x 3 with a 64 pixel overlap (50%) in both directions.

[0038] FIG. 6 illustrates a schematic workflow of the present invention, including a proof-of-concept embodiment U-net architecture. As shown, OCT data 21 is accessed / acquired, and multiple lesion characteristic images PCI are defined from the OCT data 21 as described above. For example, the lesion characteristic images PCI may include sub-RPE reflectance images, RPE internal reflectance images, retinal thickness, and / or optical attenuation coefficients (OACs), as described above with reference to FIGS. 3 and 4. Optionally, other lesion characteristic images may also be used, such as en face images / maps of specific retinal layer thickness or layer integrity, OCTA flow images (e.g., flow at or near the choriocapillaris), and / or other lesion characteristic data, as described above with reference to FIG. 5. The multiple lesion characteristic images are then combined into different (e.g., color or monochrome value) channels of a channel-encoded image 27, and the combination is then provided to a machine learning model 29, implemented herein using a U-Net architecture. The U-Net architecture is described below.

[0039] In this exemplary U-Net architecture, the convergence path consists of four convolutional neural network (CNN) blocks. Each CNN block in the convergence path may include two (e.g., 3x3) convolutions, as indicated by the asterisk symbol "*," and an activation function (e.g., a rectified linear (ReLU) unit) with optional batch normalization. The output of each CNN block in the convergence path is downsampled, such as by 2x2 max pooling, as indicated by the downward arrow. The output of the convergence path is fed to a bottleneck BN, which is shown herein to consist of two convolution layers (e.g., with batch normalization and optional 0.5 dropout). The scalability / augmentation path follows the bottleneck BN and consists herein of five CNN blocks. In the augmentation path, the output of each block is subjected to a transposed convolution (or deconvolution) to upsample (e.g., upconvert) the image / information / data. In this example, the upconversion is characterized by a 2x2 kernel (or convolution matrix), as indicated by the upward arrows. Copy-and-crop links CC1-CC4 between the corresponding downsampling and upsampling blocks copy the output of one downsampling block and connect it to the input of the corresponding upsampling block. At the end of the expansion path, the output of the last upsampling block is provided to another convolution operation (e.g., a 1x1 output convolution), as indicated by the dotted arrow, before generating its output U-out. For example, a neural network may have multiple features per pixel just before reaching the 1x1 output convolution, but the 1x1 convolution combines these multiple features into a single output value per pixel at the per-pixel level.

[0040] A combination of binary cross-entropy and Dice coefficient loss was used for training. "Icing on the cake" was used to fine-tune the model for the final layer. "Icing on the cake" is a method in which regular training is performed, and then only the final layer is (re)trained. For training, 250 macular cubes (58 with a pixel size of 512 × 128 × 1024 and 192 with a pixel size of 200 × 200 × 1024) were obtained from 155 patients using CIRRUS™ HD-OCT 4000 and 5000 (ZEISS, Dublin, CA). Experts manually drew GA outline segmentations in the en face images, focusing on both high reflectivity under the RPE and possible RPE disruption in the en face images and available B-scans. For each macular cube, a three-channel en face image was generated as described above (e.g., with reference to Figures 3 and 4A). The training and test sets of custom-generated en face images consisted of 225 eyes (187 eyes with GA, 19 eyes with drusen without GA, and 19 healthy control eyes) and 25 eyes (11 eyes with GA, 5 eyes with drusen without GA, and 9 healthy control eyes), respectively.

[0041] In processing, the trained U-Net outputs a GA segmentation 33 based on what was channel encoded in the image 21, as shown in Figure 6. The output GA segmentation 33 may be provided to a thresholding operation (and other known segmentation cleaning operations) to produce a segmentation output 35.

[0042] Segmentation by the present algorithm on a test set was compared to manual marking using qualitative and quantitative measures (e.g., area, Bland-Altman, and Pearson's correlation). Figure 7 shows the qualitative results of this proof-of-concept embodiment (e.g., the currently proposed algorithm), and Figure 8 shows the quantitative measurements of this proof-of-concept embodiment. In Figure 7, column a) shows the acquired OCT en face image, column b) shows the three-channel encoded image generated for input to the currently trained U-Net machine model, column c) shows the ground truth image (i.e., the GA regions annotated by a human expert; column d) shows the output generated by the currently proposed algorithm, and column e) shows the output from the "Advanced RPE Analysis" tool available in the current CIRRUS™ HD-OCT review software. As shown in Fig. 8, the absolute area difference and partial area difference between the GA area generated by the proposed algorithm and the GA area generated by the expert's manual marking were 0.11 ± 0.17 mm 2 and 5.51±4.7% for the advanced RPE analysis tool, compared with 0.54±0.82mm for the advanced RPE analysis tool. 2 The correlation coefficients were 25.61 ± 42.3% and 25.61 ± 42.3%. The inference time was 1183 ms per en face image using an Intel® i7 @ 2.90 GHz CPU. The correlations between the GA regions generated by the proposed algorithm and the manual markings by experts and between the GA regions generated by the advanced RPE analysis tool and the manual markings were 0.9996 (p-value < 0.001) and 0.9259 (p-value < 0.001), respectively. A Bland-Altman plot between the manual markings and the segmentation generated using the proposed algorithm showed stronger agreement than the segmentation generated using existing advanced RPE analysis tools.

[0043] 9 provides an overview of the present invention. The present method for analyzing optical coherence tomography (OCT) data to identify a particular pathology, such as GA, can begin with accessing OCT data (step S1), which may include OCTA data (or the OCTA data may be derived from the accessed OCT data). That is, the accessed data (e.g., data captured using an OCT system, data retrieved from a data store of previously captured / processed OCT data, etc.) may include OCT structural data and OCTA flow data.

[0044] Optionally, the method may further include accessing non-OCT-based data (step S2), including imaging data of an imaging modality other than OCT. For example, the system may access fundus image(s), fluorescein angiography image(s), visual field map(s), and / or non-imaging data (e.g., patient demographic data, disease and medication history, etc.).

[0045] In step S3, a set of metrics is extracted from the accessed OCT data (and optionally from other data extracted in step S2). The extracted metrics may include OCT-based metrics extracted from OCT structural data and / or OCTA-based metrics extracted from OCTA flow data. The metrics may target specific retinal layers and / or include information about the distance from the current position to predetermined retinal landmarks. For example, metrics may be extracted from each individual A-scan, and the metrics may include information about the position of the current A-scan (or axial position within the current A-scan) relative to the fovea, specific retinal layer regions, or other retinal landmarks.

[0046] In step S4, a set of images is created, each image defining or highlighting lesion-specific (e.g., GA) characteristic information. That is, the created images can characterize (e.g., be associated with) the same lesion type. The created images can be based on the extracted metrics or any of the other data types accessed in step S2. For example, the metrics extracted from each A-scan can be sorted into corresponding metric groups (e.g., in a one-to-one correspondence), and different images can be created based on each individual metric group. Images created from OCT-based data can be en face images, while images created from non-OCT-based data can be planar, frontal images. For example, the images generated may include sub-RPE reflectance, sub-RPE reflectance, en face retinal thickness, en face images of choriocapillaris flow, Sattler layer, Haller layer images, as well as fundus images (e.g., white light, red light, blue light, green light, infrared light, autofluorescence, etc.), fluorescence angiography image(s), visual field test maps, and / or 2D distributions of non-image data (e.g., patient demographic data).

[0047] In step S5, a multi-channel image is defined based on the set of images. For example, the multi-channel image can define multiple "color" channels per pixel, with each created image defining a separate color channel. In other words, the multi-channel image can include multiple image channels, each based on multiple imaging modalities. Optionally, a single color channel can be defined by the combination of the created images.

[0048] In step S6, the defined multi-channel image is provided to a machine learning model (e.g., a neural network with a U-Net architecture) trained to identify one or more lesions based on the lesion characteristic data of the individual image channels (preferably, to identify target lesions). The machine model may identify the target lesion by outlining / segmenting the lesion on the en face OCT image. That is, the individual image channel locations may be mapped to the global OCT en face image, and the identified regions of the multi-channel image where the lesion is present (based on the combination of the lesion characteristic data provided by the individual channels of each pixel of the multi-channel image) may be mapped back to the global OCT en face image.

[0049] In step S7, the identified lesions are displayed or stored in the computing device for future reference. The following describes various hardware and architectures suitable for the present invention.

[0050] visual field testing system The improvements described herein can be used in combination with any type of visual field tester / system (e.g., perimeter). One such system is the "bowl" visual field tester VF0, as shown in FIG. 10. A subject (e.g., a patient) VF1 is shown viewing a generally bowl-shaped, hemispherical projection screen (or other type of display) VF2, hence the name tester VF0. Typically, the subject is instructed to gaze at a point at the center of the hemispherical screen VF3. The subject places their head on a patient support, which may include a chin rest VF12 and / or a forehead rest VF14. For example, the subject places their head on the chin rest VF12 and their forehead on the forehead rest VF14. Optionally, the chin rest VF12 and the forehead rest VF14 can be moved together or independently of each other to properly fixate / position the patient's eyes, for example, relative to a trial lens holder VF9, which may hold a lens that allows the subject to view the screen VF2. For example, the chin rest and head rest can move independently vertically to accommodate different patient head sizes and move together horizontally and / or vertically to properly position the head, although this is not limiting and other positions / movements can be envisioned by one skilled in the art.

[0051] A projector or other image-forming device VF4, under the control of processor VF5, displays a series of test stimuli (e.g., test dots of any shape) VF6 on screen VF2. Subject VF1 indicates their viewing of stimuli VF6 by activating user input VF7 (e.g., pressing an input button). This subject's response can be recorded by processor VF5. Based on the subject's response, processor VF5 can function to evaluate the ocular field of view, for example, to determine the size, location, and / or intensity of test stimuli VF6 that are no longer visible to subject VF1, thereby determining the (visibility) threshold of test stimuli VF6. Camera VF8 can be used to capture the patient's gaze (e.g., gaze direction) throughout the test. The gaze direction can be used to confirm patient alignment and / or patient compliance with proper testing procedures. In this example, camera VF8 is positioned on the Z-axis relative to the patient's eye (e.g., relative to trial lens holder VF9) and behind the bowl (screen VF2) to capture live image(s) or video of the patient's eye. In other embodiments, this camera may be positioned away from this Z-axis. Images from gaze camera VF8 can optionally be displayed on a second display VF10 to a clinician (who may be interchangeably referred to herein as a technician) to assist with patient alignment or test verification. Camera VF8 can record and store one or more images of the eye during each stimulus presentation. This allows for the collection of tens to hundreds of images in a single visual field test, depending on the test conditions. Alternatively, camera VF8 can record and store full-length videos during the test, providing timestamps indicating when each stimulus was presented. Additionally, images can be collected between stimulus presentations to provide details of the subject's overall attention during the VF test.

[0052] A trial lens holder VF9 may be placed in front of the patient's eye to correct the eye's refractive error. Optionally, the lens holder VF9 may carry or hold a liquid trial lens that may be utilized to provide variable refractive correction to the patient VF1 (see, for example, U.S. Patent No. 8,668,338, which is incorporated herein by reference in its entirety). However, it should be noted that the present invention is not limited to using liquid trial lenses for refractive correction, and other conventional / standard trial lenses known in the art may also be used.

[0053] In some embodiments, one or more light sources (not shown) can be placed in front of the eye of subject VF1 to generate a reflection from the surface of the eye, such as the cornea. In one variation, the light source can be a light-emitting diode (LED).

[0054] While Figure 10 illustrates a projected visual field tester VF0, the invention described herein can be used with other types of devices (visual field testers), including devices that generate images via liquid crystal displays (LCDs) or other electronic displays (see, e.g., U.S. Pat. No. 8,132,916, incorporated herein by reference). Other types of visual field testers include, for example, flat screen testers, miniature testers, and binocular visual field testers. Examples of these types of testers can be found in U.S. Pat. Nos. 8,371,696, 5,912,723, 8,931,905, and U.S. Design Registration No. D472637, each of which is incorporated herein by reference in its entirety.

[0055] The visual field tester VF0 may incorporate an instrument control system (e.g., executing an algorithm, which may be software, code, and / or routines) that uses hardware signals and a motorized positioning system to automatically position the patient's eye at a desired position (e.g., the center of a refractive lens in the lens holder VF9). For example, stepper motors may move the chin rest VF12 and forehead rest VF14 under software control. Rocker switches may be provided to allow the technician to adjust the position of the patient's head by activating the chin rest and forehead stepper motors. A manually movable refractive lens may also be positioned in front of the patient's eye on the lens holder VF9 as close to the patient's eye as possible without adversely affecting patient comfort. Optionally, the instrument control algorithm may pause the performance of the visual field test while the chin rest and / or forehead motor movement is in progress if such movement would interfere with the performance of the test.

[0056] Fundus Imaging System Two categories of imaging systems used to image the fundus are flood-illumination imaging systems (or flood-illumination imagers) and scanning-illumination imaging systems (or scanning imagers). Flood-illumination imagers simultaneously flood the entire field of view (FOV) of interest of the object with light, such as with a flash lamp, and capture a full-frame image of the object (e.g., the fundus) using a full-frame camera (e.g., a camera having a two-dimensional (2D) photosensor array sized collectively to capture the desired FOV). For example, a flood-illumination fundus imager floods the fundus of the eye with light and captures a full-frame image of the fundus in a single image capture sequence of the camera. Scanning imagers provide a scanning beam that is scanned across the object, e.g., the eye, and the scanning beam is imaged at different scanning locations as the scanning beam is scanned across the object, creating a series of image segments that can be reconfigured, e.g., combined, to create a composite image of the desired FOV. The scanning beam can be a point, a line, or a two-dimensional region such as a slit or a wide line. Examples of fundus imaging devices are described in US Pat. Nos. 8,967,806 and 8,998,411.

[0057] FIG. 11 illustrates an example of a slit-scanning ophthalmic system SLO-1 for imaging the fundus F, which is the inner surface of the eye E opposite the ocular lens (or crystalline lens) CL and may include the retina, optic disc, macula, fovea, and posterior pole. In this example, the imaging system is in a so-called “scan-descan” configuration, in which a scanning line beam SB traverses the optical components of the eye E (including the cornea Crn, iris Irs, pupil Ppl, and lens) to scan across the entire fundus F. In the case of a projected light fundus imaging device, no scanner is required, and light is irradiated across the entire desired field of view (FOV) at once. Other scanning configurations are known in the art, and the specific scanning configuration is not critical to the present invention. As shown, the imaging system includes one or more light sources LtSrc, preferably a polychromatic LED system or a laser system with a suitably adjusted etendue. An optional slit Slt (adjustable or stationary) may be positioned in front of the light source LtSrc and used to adjust the width of the scanning line beam SB. Additionally, the slit Slt can remain stationary during imaging or can be adjusted to different widths to allow for different confocal levels and different applications for a particular scan or during scanning used for reflection suppression. An optional objective lens ObjL can be placed before the slit Slt. The objective lens ObjL can be any one of several state-of-the-art lenses, including, but not limited to, refractive, diffractive, reflective, or hybrid lenses / systems. Light from the slit Slt passes through a pupil-splitting mirror SM and is directed to the scanner LnScn. It is desirable to keep the scan plane and pupil plane as close as possible to reduce vignetting in the system. An optional optical system DL can be included to manipulate the optical distance between the images of the two components. The pupil-splitting mirror SM can pass the illumination beam from the light source LtSrc to the scanner LnScn and reflect the detection beam from the scanner LnScn (e.g., reflected light returning from the eye E) toward the camera Cmr. The task of the pupil-splitting mirror SM is to split the illumination beam from the scanner LnScn and assist in suppressing system reflections.Scanner LnScn can be a rotating galvo scanner or other type of scanner (e.g., piezoelectric or voice coil, microelectromechanical system (MEMS) scanner, electro-optic deflector, and / or rotating polygon scanner). Depending on whether pupil splitting occurs before or after scanner LnScn, the scan can be split into two steps with one scanner in the illumination path and a separate scanner in the detection path. Specific pupil splitting arrangements are described in detail in U.S. Pat. No. 9,456,746, the entire contents of which are incorporated herein by reference.

[0058] From the scanner LnScn, the illumination beam passes through one or more optical systems—in this case, a scan lens SL and an ophthalmic or ocular lens OL—that enable the pupil of the eye E to be imaged onto the system's image pupil. Generally, the scan lens SL receives the scanning illumination beam from the scanner LnScn at any of a number of scan angles (angles of incidence) and generates a scan line beam SB having a substantially planar focal plane (e.g., a collimated optical path). The ophthalmic lens OL can focus the scan line beam SB onto an imaging object. In this example, the ophthalmic lens OL can focus the scan beam SB onto the fundus F (or retina) of the eye E to image the fundus. In this way, the scan line beam SB creates a transverse scan line that moves across the fundus F. One possible configuration of these optical systems is a Keplerian telescope, in which the distance between the two lenses is selected to generate a nearly telecentric intermediate fundus image (4-f configuration). The ophthalmic lens OL can be a single lens, an achromatic lens, or an arrangement of different lenses. All lenses can be refractive, diffractive, reflective, or hybrid, as known to those skilled in the art. The size and / or shape of the ophthalmic lens OL, the scanning lens SL, the pupil-splitting mirror SM, and the scanner LnScn can vary depending on the desired field of view (FOV). Therefore, an arrangement can be envisioned in which multiple components can be switched in and out of the beam path, for example, by using a flip of the optical system, a motorized wheel, or removable optical elements, depending on the field of view. Because a change in field of view results in a different beam size on the pupil, the pupil splitting can also be changed to accommodate a change in FOV. For example, a field of view of 45° to 60° is a typical or standard FOV for a fundus camera. Higher fields of view, such as wide-field FOVs of 60° to 120° or more, may also be feasible. A wide-field FOV may be desirable for combining a wide-line fundus imager (BLFI) with another imaging modality, such as optical coherence tomography (OCT). The upper limit of the field of view may be determined by the accessible working distance combined with the physiological conditions around the human eye. Since a typical human retina has an FOV of 140° horizontally and 80°-100° vertically, it may be desirable to have an asymmetric field of view with as high an FVO as possible on the system.

[0059] The scanning beam SB passes through the pupil Ppl of the eye E and is directed onto the retina or fundus, i.e., surface F. The scanner LnScn1 adjusts the position of the light on the retina or fundus F so that a range of lateral positions of the eye E is illuminated. The reflected or scattered light (or emitted light in the case of fluorescence imaging) is directed along a path similar to the illumination, defining a focused beam CB on a detection path to the camera Cmr.

[0060] In the “scan-descan” configuration of the exemplary slit-scanning ophthalmic system SLO-1 of the present invention, light returning from eye E is “descanned” by scanner LnScn on its way to pupil-splitting mirror SM. That is, scanner LnScn scans illumination beam SB from pupil-splitting mirror SM to define a scanning illumination beam SB across eye E, but because scanner LnScn also receives returning light from eye E at the same scanning position, it effectively descans the returning light (e.g., cancels the scanning motion) to define a non-scanning (e.g., stationary) focused beam from scanner LnScn to pupil-splitting mirror SM, which folds the focused beam toward camera Cmr. At pupil-splitting mirror SM, reflected light (or emitted light, in the case of fluorescence imaging) is separated from the illumination light onto a detection path that is directed to camera Cmr, which may be a digital camera with a photosensor for capturing an image. An imaging (e.g., objective) lens ImgL may be positioned in the detection path so that the fundus is imaged onto camera Cmr. As with the objective lens ObjL, the imaging lens ImgL can be any type of lens known in the art (e.g., refractive, diffractive, reflective, or hybrid lens). Additional operational details, particularly methods for reducing artifacts in images, are described in International Publication WO 2016 / 124644, the entire contents of which are incorporated herein by reference. The camera Cmr captures the received images and, for example, creates an image file, which can be further processed by one or more (electronic) processors or computing devices (e.g., the computer system of FIG. 20). Thus, focused beams (returned from all scanning positions of the scanning line beam SB) are collected by the camera Cmr, and a full-frame image Img can be constructed, such as by montage, from a combination of the individually captured focused beams. However, other scanning configurations are also envisioned, including those in which the illumination beam is scanned across the eye E and the focused beam is scanned across the camera's photosensor array.WO 2012 / 059236 and U.S. Patent Application Publication No. 2015 / 0131050, which are incorporated herein by reference, describe several embodiments of scanning slit ophthalmoscopes, including various designs, such as designs in which the returning light is swept across the camera's photosensor array, and designs in which the returning light is not swept across the camera's photosensor array.

[0061] In this example, the camera Cmr is connected to a processor (e.g., processing module) Proc and a display (e.g., display module, computer screen, electronic screen, etc.) Dspl, both of which may be part of the imaging system itself or may be part of separate, dedicated processing and / or display units, such as a computer system, with data passed from the camera Cmr to the computer system via a cable or computer network, including a wireless network. The display and processor may be an integrated unit. The display may be a traditional electronic display / screen or may be a touchscreen and may include a user interface for displaying information to and receiving information from an equipment operator or user. A user may interact with the display using any type of user input device known in the art, including, but not limited to, a mouse, knob, button, pointer, and touchscreen.

[0062] It may be desirable for the patient's gaze to remain fixed while imaging is being performed. One way to achieve gaze fixation is to provide a fixation target to which the patient can be instructed to gaze. The fixation target can be internal or external to the device, depending on which region of the eye is being imaged. One embodiment of an internal fixation target is shown in FIG. 11. In addition to the primary light source LtSrc used for imaging, an optional second light source FxLtSrc, such as one or more LEDs, can be positioned to image a light pattern onto the retina using a lens FxL, scanning elements FxScn, and a reflector / mirror FxM. The fixation scanner FxScn can move the position of the light pattern, and the reflector FxM directs the light pattern from the fixation scanner FxScn to the fundus F of the eye E. Preferably, the fixation scanner FxScn is positioned at the pupil plane of the system so that the light pattern on the retina / fundus can be moved according to the desired fixation position.

[0063] Slit-scanning ophthalmoscope systems can operate in different imaging modes depending on the light source and wavelength-selective filtering elements used. True-color reflectance imaging (similar to that observed by clinicians when examining the eye using a handheld or slit-lamp ophthalmoscope) can be achieved when imaging the eye using a series of colored LEDs (red, blue, and green). Images for each color can be built up stepwise with each LED turned on at each scanning position, or each color image can be captured completely separately. The three color images can be combined to display a true-color image, or displayed individually to highlight different features of the retina. The red channel best highlights the choroid, the green channel highlights the retina, and the blue channel highlights the pre-retinal layers. Additionally, specific frequencies of light (e.g., individual colored LEDs or lasers) can be used to excite different fluorophores (e.g., autofluorescence) within the eye, and the resulting fluorescence can be detected by filtering out the excitation wavelengths.

[0064] Fundus imaging systems can also provide infrared reflectance images, such as by using an infrared laser (or other infrared light source). Infrared (IR) mode is advantageous in that the eye is not sensitive to IR wavelengths. This infrared (IR) mode may allow the user to continuously capture images without obstructing the eye (e.g., in preview / alignment mode) to assist the user during device alignment. IR wavelengths also have high penetration through tissue and may provide improved visualization of choroidal structures. Additionally, fluorescein angiography (FA) and indocyanine green (ICG) angiography imaging can be achieved by collecting images after a fluorescent dye is injected into the subject's bloodstream. For example, with FA (and / or ICG), a series of time-lapse images may be captured after injecting a photoreactive dye (e.g., a fluorescent dye) into the subject's bloodstream. Note that fluorescent dyes can cause life-threatening allergic reactions in some individuals, so caution is advised. High-contrast grayscale images are captured using specific light frequencies selected to excite the dye. As the dye flows through the eye, various parts of the eye glow brightly (e.g., fluoresce), allowing the viewer to see how the dye, and therefore blood, is progressing through the eye.

[0065] Optical coherence tomography system Generally, optical coherence tomography (OCT) uses low-coherence light to generate two-dimensional (2D) and three-dimensional (3D) internal views of biological tissues. OCT enables in vivo imaging of retinal structures. OCT angiography (OCTA) generates flow information, such as vascular flow, from within the retina. Examples of OCT systems are provided in U.S. Patent Nos. 6,741,359 and 9,706,915, and examples of OCTA systems are provided in U.S. Patent Nos. 9,700,206 and 9,759,544, all of which are incorporated herein by reference in their entireties. An exemplary OCT / OCTA system is provided herein.

[0066] FIG. 12 illustrates a generalized frequency-domain optical coherence tomography (FD-OCT) system for collecting 3D image data of the eye suitable for use with the present invention. The FD-OCT system OCT_1 includes a light source LtSrc1. Typical light sources include, but are not limited to, a broadband light source with a short temporal coherence length or a swept laser source. A beam of light from the light source LtSrc1 is typically guided by an optical fiber Fbr1 to illuminate a sample, such as an eye E, a typical sample being human intraocular tissue. The light source LrSrc1 may be, for example, a broadband light source with a short temporal coherence length in the case of spectral-domain OCT (SD-OCT) or a tunable laser source in the case of swept-source OCT (SS-OCT). The light may typically be scanned using a scanner Scnr1 between the output of the optical fiber Fbr1 and the sample E, such that the beam of light (dashed line Bm) is scanned laterally across the region of the sample to be imaged. The light beam from scanner Scnr1 passes through a scan lens SL and an ophthalmic lens OL and can be focused onto the sample E to be imaged. The scan lens SL can receive the light beam from scanner Scnr1 at multiple angles of incidence to generate substantially collimated light, which the ophthalmic lens OL can then focus onto the sample. This example shows a scanning beam that needs to be scanned in two lateral directions (e.g., the x and y directions on a Cartesian plane) to scan a desired field of view (FOV). This example is point-field OCT, which uses a point-field beam to scan across the sample. Thus, scanner Scnr1 is illustratively shown to include two sub-scanners: a first sub-scanner Xscn for scanning the point-field beam across the sample in a first direction (e.g., the horizontal x direction) and a second sub-scanner Yscn for scanning the point-field beam across the sample in an intersecting second direction (e.g., the vertical y direction). If the scanning beam is a line field beam (e.g., line field OCT) and can sample an entire line portion of the sample at a time, only one scanner may be required to scan the line field beam across the sample to span the desired FOV.If the scanning beam is a full-field beam (eg, full-field OCT), a scanner may not be required and the full-field light beam may be illuminated across the entire desired FOV at once.

[0067] Regardless of the type of beam used, light scattered from the sample (e.g., sample light) is collected. In this example, scattered light returning from the sample is collected into the same optical fiber Fbr1 used to route light for illumination. A reference beam derived from the same light source LtSrc1 travels along a separate path, which in this case includes optical fiber Fbr2 and a retroreflector RR1 with an adjustable optical delay. As will be appreciated by those skilled in the art, a transmissive reference path can also be used, with an adjustable delay located in either the sample or the reference arm of the interferometer. The collected sample light is combined with the reference beam, for example, in a fiber coupler Cplr1, to form optical interference within an OCT photodetector Dtctr1 (e.g., a photodetector array, digital camera, etc.). While one fiber port is shown reaching detector Dtctr1, those skilled in the art will appreciate that various interferometer designs can be used for balanced or unbalanced detection of the interference signal. The output from detector Dtctr1 is fed to a processor (e.g., an internal or external computing device) Cmp1, which converts the observed interference into sample depth information. The depth information may be stored in a memory associated with the processor Cmp1 and / or displayed on a display (e.g., computer / electronic display / screen) Scn1. The processing and storage functions may be localized within the OCT device, or the functions may be offloaded to (e.g., executed on) an external processor (e.g., an external computer system) to which the collected data is transferred. An example of a computing device (or computer system) is shown in Figure 20. This unit may be dedicated to data processing or may perform other tasks that are quite general and not dedicated to the OCT device.The processor (computing device) Cmp1 may include, for example, a field programmable gate array (FPGA), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a graphics processing unit (GPU), a system on a chip (SoC), a central processing unit (CPU), a general purpose graphics processing unit (GPGPU), or combinations thereof, which may perform some or all of the processing steps in a serial and / or parallel manner with one or more host processors and / or one or more external computing devices.

[0068] The sample and reference arms in an interferometer can be constructed of bulk optics, fiber optics, or hybrid bulk optics systems and can have different architectures, such as Michelson, Mach-Zehnder, or common-path designs, as known to those skilled in the art. Light beam, as used herein, should be interpreted as any carefully directed optical path. Instead of mechanically scanning the beam, a light field can illuminate a one-dimensional or two-dimensional area of ​​the retina to generate OCT data (e.g., U.S. Pat. No. 9,332,902; D. Hillmann et al., "Holoscopy-holographic optical coherence tomography," Optics Letters, Vol. 36(13), p. 2290, 2011; Y. Nakamura et al., "High-Speed ​​three dimensional human retinal imaging by line field spectral domain optical coherence tomography," Optics Express, 2011). Express, 15(12), p. 7103, 2007; Blazkiewicz et al., "Signal-to-noise ratio study of full-field Fourier-domain optical coherence tomography," Applied Optics, 44(36), p. 7722 (2005). In time-domain systems, the reference arm must have an adjustable optical delay to create interference. Balanced detection systems are typically used in TD-OCT and SS-OCT systems, while a spectrometer is used at the detection port for SD-OCT systems. The invention described herein can be applied to either type of OCT system.Various aspects of the present invention may be applied to any type of OCT system or other types of ophthalmic diagnostic systems and / or multiple ophthalmic diagnostic systems, including, but not limited to, fundus imaging systems, visual field testing devices, and scanning laser polarimeters.

[0069] In Fourier-domain optical coherence tomography (FD-OCT), each measurement is a real-valued spectrally controlled interferogram (Sj(k)). The real-valued spectral data typically undergoes several post-processing steps, including background subtraction, dispersion correction, etc. A Fourier transform of the processed interferogram yields a complex OCT signal output Aj(z) = |Aj|eiφ. The absolute value of this complex OCT signal, |Aj|, ​​reveals the scattering intensity at different path lengths and, therefore, the scattering profile with respect to depth (z-direction) within the sample. Similarly, the phase φj can also be extracted from the complex OCT signal. The scattering profile with respect to depth is called an axial scan (A-scan). A collection of A-scans measured at adjacent locations within the sample produces a cross-sectional image (tomogram or B-scan) of the sample. A collection of B-scans collected at different lateral locations on the sample constitutes a data volume or cube. For a particular data volume, the fast axis refers to the scan direction along one B-scan, and the slow axis refers to the axis along which multiple B-scans are collected. The term "cluster scan" may refer to a unit or block of data generated by repeated acquisition at the same (or substantially the same) location (or region) to analyze motion contrast, which may be used to identify blood flow. A cluster scan can consist of multiple A-scans or B-scans collected at approximately the same location on the sample at a relatively short time interval. Because the scans in a cluster scan are of the same region, stationary structures remain relatively unchanged between scans in the cluster scan, while motion contrast between scans that meet predetermined criteria may be identified as blood flow.

[0070] Various methods for generating B-scans are known in the art, including, but not limited to, along the horizontal or x-direction, along the vertical or y-direction, along the x and y diagonals, or in a circular or spiral pattern. A B-scan may be in the xz dimension, but may also be any cross-sectional image including the z-dimension. An exemplary OCT B-scan image of a normal retina of a human eye is shown in FIG. 13. An OCT B-scan of the retina provides a view of the structure of the retinal tissue. For illustrative purposes, FIG. 13 identifies the various normal retinal layers and layer boundaries. The identified retinal boundary layers include (from top to bottom) the inner limiting membrane (ILM) layer 1, the retinal nerve fiber layer (RNFL or NFL) layer 2, the ganglion cell layer (GCL) layer 3, the inner plexiform layer (IPL) layer 4, the inner nuclear layer (INL) layer 5, the outer plexiform layer (OPL) layer 6, the outer nuclear layer (ONL) layer 7, the junction between the outer segments (OS) and inner segments (IS) of photoreceptors (indicated by reference numeral 8), the external limiting membrane (ELM or OLM) layer 9, the retinal pigment epithelium (RPE) layer 10, and the Bruch's membrane (BM) layer 11.

[0071] In OCT angiography or functional OCT, analysis algorithms may be applied to OCT data collected at different times (e.g., cluster scans) at the same or nearly the same sample location on the sample to analyze motion or flow (see, e.g., U.S. Patent Application Publication Nos. 2005 / 0171438, 2012 / 0307014, 2010 / 0027857, 2012 / 0277579, and U.S. Patent No. 6,549,801, all of which are incorporated by reference in their entireties). OCT systems may use any one of a number of OCT angiography processing algorithms (e.g., motion contrast algorithms) to identify blood flow. For example, motion contrast algorithms can be applied to intensity information derived from the image data (intensity-based algorithms), phase information from the image data (phase-based algorithms), or complex image data (complex-based algorithms). An en face image is a 2D projection of the 3D OCT data (e.g., by averaging the intensity of each individual A-scan, whereby each A-scan defines a pixel in the 2D projection). Similarly, an en face vascular image is an image that displays motion contrast signals in which the data dimension corresponding to depth (e.g., the z-direction along the A-scan) is displayed as a single representative value (e.g., a pixel in the 2D projection image), typically by summing or integrating all or isolated portions of the data (see, e.g., U.S. Pat. No. 7,301,644, incorporated herein by reference in its entirety). OCT systems that provide angiography capabilities may be referred to as OCT angiography (OCTA) systems.

[0072] FIG. 14 shows an example of an en face vasculature image. After processing the data and highlighting motion contrast using any of the motion contrast methods known in the art, an en face (e.g., front view) image of the vasculature may be generated by summing pixel ranges corresponding to a tissue depth from the surface of the retinal internal limiting membrane (ILM). FIG. 15 shows an exemplary B-scan of a vasculature (OCTA) image. As shown, structural information may be less clear because blood flow traverses multiple retinal layers, obscuring them more than in a structural OCT B-scan such as that shown in FIG. 13. Nevertheless, OCTA provides a noninvasive technique for imaging the retinal and choroidal microvasculature, which may be important for diagnosing and / or monitoring various pathologies. For example, OCTA may be used to identify diabetic retinopathy by identifying microaneurysms, neovascular complexes, and quantifying the foveal avascular zone and nonperfused areas. Furthermore, OCTA has been shown to show good agreement with fluorescein angiography (FA), a more traditional but less invasive technique that requires the injection of dye to observe vascular flow in the retina. Furthermore, in dry age-related macular degeneration (AMD), OCTA has been used to monitor the overall decrease in choriocapillaris flow. Similarly, in exudative AMD, OCTA can provide qualitative and quantitative analysis of choroidal neovascular membranes. OCTA has also been used to study vascular obstruction, for example, to assess nonperfused areas and the integrity of the superficial and deep plexuses.

[0073] Neural Networks As mentioned above, the present invention may use neural network (NN) machine learning (ML) models. For completeness, neural networks are generally described herein. The invention may use any of the following neural network architectures, alone or in combination: A neural network, or neural net, is a network of interconnected neurons (via nodes), with each neuron representing a node in the network. Collections of neurons may be arranged in layers, with the output of one layer being fed forward to the next layer in a multi-layer perceptron (MLP) arrangement. An MLP may be understood as a feed-forward neural network that maps a set of input data to a set of output data.

[0074] FIG. 16 illustrates an example of a multilayer perceptron (MLP) neural network. The structure may include multiple hidden (e.g., inner) layers HL1 through HLn, which map an input layer InL (which receives a set of inputs (or vector inputs) in_1 through in_3) to an output layer OutL, which generates a set of outputs (or vector outputs), e.g., out_1 and out_2. Each layer may have any number of nodes, which are shown here as circles within each layer for illustrative purposes. In this example, the first hidden layer HL1 has two nodes, and hidden layers HL2, HL3, and HLn each have three nodes. Generally, the deeper the MLP (e.g., the greater the number of hidden layers in the MLP), the greater its learning capacity. The input layer InL may receive vector inputs (shown for illustrative purposes as a three-dimensional vector consisting of in_1, in_2, and in_3) and feed the received vector inputs to the first hidden layer HL1 in the sequence of hidden layers. The output layer OutL receives the output from the last hidden layer in the multi-layer model, say HLn, and produces a vector output result (shown for illustration purposes as a two-dimensional vector consisting of out_1 and out_2).

[0075] Typically, each neuron (i.e., node) generates one output, which is fed forward to neurons in the immediately succeeding layer. However, each neuron in a hidden layer may receive multiple inputs, either from the input layer or from the outputs of neurons in the immediately preceding hidden layer. In general, each node may apply a function to its inputs to generate the output for that node. Nodes in a hidden layer (e.g., the training layer) may apply the same function to each of their inputs to generate their respective outputs. However, some nodes, e.g., nodes in the input layer InL, may receive only one input and be passive, meaning that they simply relay the value of that one input to their output, e.g., they provide a copy of that input to their output; this is indicated by the dashed arrows in the nodes of the input layer InL for illustrative purposes.

[0076] For illustrative purposes, FIG. 17 shows a simplified neural network consisting of an input layer InL′, a hidden layer HL1′, and an output layer OutL′. The input layer InL′ is shown to have two input nodes i1 and i2, which receive inputs Input_1 and Input_2, respectively (e.g., the input nodes of layer InL′ receive a two-dimensional input vector). The input layer InL′ feeds forward into one hidden layer HL1′ with two nodes h1 and h2, which in turn feeds forward into an output layer OutL′ with two nodes o1 and o2. The interconnections, or links, between neurons (shown with solid arrows for illustrative purposes) have weights w1 through w8. Typically, except for the input layer, a node (neuron) may receive as input the output of a node in the previous layer. Each node may calculate its output by multiplying each of its inputs by each input's corresponding interconnection weight, summing the products of the inputs, adding (or multiplying) a constant defined by other weights or biases that may be associated with that particular node (e.g., node weights w9, w10, w11, and w12 corresponding to nodes h1, h2, o1, and o2, respectively), and then applying a nonlinear or logarithmic function to the result. The nonlinear function may be referred to as an activation function or a transfer function. Multiple activation functions are known in the art, and the selection of a particular activation function is not important to this description. However, it should be noted that the operation of an ML model, the behavior of a neural net, depends on the values ​​of the weights, which the neural network may be trained to provide a desired output for a given input.

[0077] During a training, or learning, phase, a neural network learns (e.g., is trained to identify) appropriate weight values ​​to achieve a desired output for a given input. Before a neural network is trained, each weight may be individually assigned an initial (e.g., random, optionally non-zero) value, such as a random number seed. Various methods for assigning initial weights are known in the art. The weights are then trained (optimized) so that, for a given training vector input, the neural network produces an output that approximates a desired (predetermined) training vector output. For example, the weights may be gradually adjusted over thousands of iterative cycles by a method called backpropagation. In each backpropagation cycle, a training input (e.g., a vector input or training input image / sample) is passed forward through the neural network to provide its actual output (e.g., a vector output). The error of each output neuron, or output node, is then calculated based on the actual neuron's output and the supervised training output for that neuron (e.g., a training output image / sample corresponding to the current training input image / sample). It then propagates backward through the neural network (from the output layer back to the input layer), updating the weights based on how much influence each weight has on the overall error, thereby moving the neural network's output closer to the desired training output. This cycle is then repeated until the neural network's actual output is within an acceptable error range of the desired training output for that training input. As will be appreciated, each training input may require many backpropagation iterations to achieve the desired error range. Typically, an epoch refers to one backpropagation iteration (e.g., one forward pass and one backward pass) of all training samples, and training a neural network may require many epochs. Generally, the larger the training set, the better the performance of the trained ML model; therefore, various data augmentation methods may be used to increase the size of the training set.For example, if the training set includes pairs of corresponding training input images and training output images, the training images may be divided into multiple corresponding image segments (or patches). Corresponding patches from the training input images and training output images may be paired to define multiple training patch pairs from one input / output image pair, thereby expanding the training set. However, training a large training set increases the demands on computer resources, such as memory and data processing resources. The computational demands may be reduced by dividing the large training set into multiple mini-batches, the size of which determines the number of training samples in one forward / backward pass. In this case, one epoch may contain multiple mini-batches. Another problem is the possibility that the neural network may overfit the training set, reducing its ability to generalize from a specific input to different inputs. The overfitting problem may be mitigated by creating an ensemble of neural networks or by randomly dropping out nodes in the neural network during training, which effectively removes the dropped leads from the neural network. Various dropout adjustment methods, such as inverse dropout, are known in the art.

[0078] It should be noted that the operation of a trained NN machine model is not a simple algorithm of computation / analysis steps. Indeed, when a trained NN machine model receives an input, the input is not analyzed in the traditional sense. Rather, regardless of the purpose or nature of the input (e.g., vectors defining a live image / scan or vectors defining any other entity such as a demographic description or activity record), the input is subjected to the same architectural construction of the trained neural network (e.g., the same node / layer arrangement, trained weights and bias values, predetermined convolution / deconvolution operations, activation functions, pooling operations, etc.), and it may not be obvious how the architectural construction of the trained network generates its output. Furthermore, the values ​​of the trained weights and biases are not deterministic and depend on many factors, such as the amount of time given to the neural network for training (e.g., the number of epochs in training), the random starting values ​​of the weights before training begins, the computer architecture of the machine on which the NN is trained, the selection of training samples, the distribution of training samples among multiple mini-batches, the selection of activation functions, the selection of error functions that modify the weights, and even whether training is interrupted on one machine (e.g., with a first computer architecture) and completed on another machine (e.g., with a different computer architecture). The point is that the reasons why a trained ML model arrives at a particular output are not clear, and much research is currently being conducted to identify the factors on which ML models base their output. Therefore, the processing of neural networks on live data cannot be reduced to a simple algorithmic step. Rather, the operation depends on the training architecture, training sample set, training sequence, and various circumstances in the training of the ML model.

[0079] In general, constructing a neural network machine learning model may include a learning (or training) stage and a classification (or computation) stage. In the learning stage, a neural network may be trained for a specific purpose and provided with a set of training examples, including training (sample) inputs and training (sample) outputs, and optionally a set of validation examples for testing the training progress. During this learning process, various weights associated with nodes and node interconnections within the neural network are gradually adjusted to reduce the error between the neural network's actual output and the desired training output. In this way, a multi-layer feedforward neural network (such as those described above) may be able to approximate any measurable function to any desired accuracy. The result of the learning stage is a learned (e.g., trained) (neural network) machine learning (ML). In the computation stage, a set of test inputs (or live inputs) may be provided to the learned (trained) ML model, which may apply what it has learned to generate output predictions based on the test inputs.

[0080] Similar to the conventional neural networks of Figures 16 and 17, convolutional neural networks (CNNs) are also composed of neurons with learnable weights and biases. Each neuron receives an input and performs an operation (e.g., a dot product), optionally followed by a nonlinear transformation. However, CNNs receive raw image pixels at one end (e.g., the input) and provide a classification (or class) score at the other end (e.g., the output). Because CNNs expect images as input, they are optimized to handle volumes (e.g., image pixel height and width, and image depth, e.g., color depth, such as RGB depth defined by three colors: red, green, and blue). For example, CNN layers may be optimized for neurons arranged in three dimensions. Neurons in a CNN layer may connect to a small region of the previous layer rather than all of the neurons in a fully connected NN. The final output layer of a CNN may reduce the full image to a single vector (classification) arranged along the depth dimension.

[0081] FIG. 18 provides an exemplary convolutional neural network architecture. A convolutional neural network may be defined as a sequence of two or more layers (e.g., Layer 1 through Layer N), where a layer may include an (image) convolution step, a (resulting) weighted sum step, and a nonlinear function step. The convolution may be performed on the input data by, for example, applying a filter (or kernel) on a moving window over the input data to generate a feature map. Each layer and layer component may have different predetermined filters (from a filter bank), weights (or weighting parameters), and / or function parameters. In this example, the input data may be an image of a certain pixel height and width, or the raw pixel values ​​of this image. In this example, the input image is depicted as having a depth of three color channels, RGB (red, green, blue). Optionally, various preprocessing steps may be performed on the input image, and the results of the preprocessing steps may be input instead of or in addition to the raw image data. Some examples of image processing may include retinal vessel map segmentation, color space conversion, adaptive histogram equalization, connected component generation, etc. Within a layer, a dot product may be calculated between certain weights and their connected small regions within the input volume. While many methods for constructing CNNs are known in the art, by way of example, layers may be configured to apply element-wise activation functions, such as a max(0,x) threshold at zero. Pooling functions may be performed (e.g., along the x and y directions) to downsample the volume. Fully connected layers may be used to identify classification outputs and generate one-dimensional output vectors, which have proven useful in image recognition and classification. However, for image segmentation, CNNs must classify each pixel. Because each CNN layer tends to reduce the resolution of the input image, another stage is required to upsample the image to its original resolution. This may be achieved by applying a transposed convolution (or deconvolution) stage TC, which typically does not use any predetermined interpolation method but instead has learnable parameters.

[0082] Convolutional neural networks have been successfully applied to many problems in computer vision. As mentioned above, training a CNN generally requires a large training dataset. The U-Net architecture is based on a CNN and can generally be trained with a smaller training dataset than traditional CNNs.

[0083] FIG. 19 illustrates an exemplary U-Net architecture. This exemplary U-Net includes an input module (or input layer or stage), which receives an input U-in (e.g., an input image or image patch) of any size. For convenience, the image size at any stage or layer is indicated within a box representing the image; for example, in the input module, the number "128x128" is enclosed, indicating that the input image U-in is composed of 128x128 pixels. The input image may be a fundus image, an OCT / OCTA en face image, a B-scan image, etc. However, it should be understood that the input may be of any size or dimensionality. For example, the input image may be a multi-channel image (RGB color image), a monochrome image, a volumetric image, etc. The input image passes through a series of processing layers, each of which is illustrated with exemplary sizes, but these sizes are for illustrative purposes only and will depend, for example, on the size of the image, the convolutional filters, and / or the pooling stage. This architecture consists of a convergent path (exemplary herein including four encoding modules) followed by an augmented path (exemplary herein including four decoding modules), with copy-and-crop links (e.g., CC1-CC4) between corresponding modules / stages that copy the output of one encoding module in the convergent path and connect (e.g., append) it to the upconverted input of the corresponding decoding module in the augmented path. This results in a characteristic U-shape, from which the architecture is named. Optionally, for computational considerations, a "bottleneck" module / stage (BN) can be placed between the convergent path and the augmented path. The bottleneck BN may consist of two convolutional layers (with batch normalization and optional dropout).

[0084] The convergent path is similar to an encoder and typically uses feature maps to capture context (or feature) information. In this example, each encoding module in the convergent path includes two or more convolutional layers, indicated by an asterisk symbol "*," which may be followed by a max-pooling layer (e.g., a downsampling layer). For example, the input image U-in is shown passing through two convolutional layers, each with 32 feature maps. It can be understood that each convolutional kernel produces a feature map (e.g., the output from a convolution operation with a given kernel is an image commonly referred to as a "feature map"). For example, the input U-in passes through a first convolution that applies 32 convolutional kernels (not shown), producing an output consisting of 32 individual feature maps. However, as is known in the art, the number of feature maps produced by a convolution operation can be adjusted (upward or downward). For example, the number of feature maps can be reduced by averaging groups of feature maps, deleting some feature maps, or other known methods of reducing feature maps. In this example, this first convolution is followed by a second convolution whose output is limited to 32 feature maps. Another way to envision the feature maps is to consider the output of the convolution layer as a 3D image whose 2D dimensions are given by the stated XY plane pixel dimensions (e.g., 128x128 pixels) and whose depth is given by the number of feature maps (e.g., the depth of the 32 plane images). Following this illustration, the output of the second convolution (e.g., the output of the first encoding module in the convergence path) can be described as a 128x128x32 image. The output from the second convolution is then subjected to a pooling operation, which reduces the 2D dimensions of each feature map (e.g., the X and Y dimensions can each be reduced by half). The pooling operation can be embodied within a downsampling process, as indicated by the downward arrows. Several pooling methods, such as max pooling, are known in the art, and the particular pooling method is not critical to the present invention.The number of feature maps doubles with each pooling, such as 32 feature maps in the first encoding module (or block), 64 feature maps in the second encoding module, and so on. Thus, the convergence path forms a convolutional network composed of multiple encoding modules (or stages or blocks). As is typical for convolutional networks, each encoding module may provide at least one convolution stage followed by an activation function (e.g., a rectified linear unit (ReLU) or sigmoid layer) (not shown) and a max-pooling operation. Generally, the activation function introduces nonlinearity into the layer (e.g., to avoid overfitting), receives the layer's results, and determines whether to "activate" the output (e.g., whether the value at a particular node meets a predetermined criterion to forward the output to the next layer / node). In summary, the convergence path generally reduces spatial information and increases feature information.

[0085] The extension path is similar to the decoder, notably providing localization and spatial information to the results of the convergence path, despite the downsampling and any max-pooling performed in the contraction stage. The extension path includes multiple decoding modules, each of which combines its current upconverted input with the output of a corresponding encoding module. Thus, features and spatial information are combined in the extension path through a series of upconvolutions (e.g., upsampling or transposed convolutions, i.e., deconvolutions) and combinations (e.g., via CC1-CC4) with high-resolution features from the convergence path. Thus, the output of the deconvolution layer is combined with the corresponding (optionally cropped) feature map from the convergence path, followed by two convolutional layers and activation functions (optionally batch normalized).

[0086] The output from the last augmentation module in the augmentation path may be fed to another processing / training block or layer, such as a classifier block, which may be trained with the U-Net architecture. Alternatively, or additionally, the output of the last upsampling block (at the end of the augmentation path) may be provided to another convolution (e.g., output convolution) operation, as indicated by the dotted arrow, before generating its output U-out. The kernel size of the output convolution may be selected to reduce the dimensions of the last upsampling block to a desired size. For example, a neural network may have multiple features per pixel just before reaching the output convolution, which may provide a 1×1 convolution operation to combine these multiple features into a single output value per pixel at the per-pixel level.

[0087] Computing Devices / Systems FIG. 20 illustrates an exemplary computer system (or computing device). In some embodiments, one or more computer systems may provide functionality described or illustrated herein and / or perform one or more steps of one or more methods described or illustrated herein. The computer system may take any suitable physical form. For example, the computer system may be an embedded computer system, a system-on-chip (SOC), or a single-board computer system (SBC) (e.g., a computer-on-module (COM) or system-on-module (SOM)), a desktop computer system, a laptop or notebook computer system, a mesh of computer systems, a mobile phone, a personal digital assistant (PDA), a server, a tablet computer system, an augmented / virtual reality device, or a combination of two or more of these. Where appropriate, the computer system may reside in a cloud, which may include one or more cloud components within one or more networks.

[0088] In some embodiments, a computer system may include a processor Cpnt1, a memory Cpnt2, a storage Cpnt3, an input / output (I / O) interface Cpnt4, a communication interface Cpnt5, and a bus Cpnt6. The computer system may also optionally include a display Cpnt7, such as a computer monitor or screen.

[0089] The processor Cpnt1 includes hardware for executing instructions, such as those that constitute a computer program. For example, the processor Cpnt1 may be a central processing unit (CPU) or a general-purpose computing-on-graphics processing unit (GPGPU). The processor Cpnt1 may read (or fetch) instructions from an internal register, an internal cache, memory Cpnt2, or storage Cpnt3, decode and execute the instructions, and write one or more results to the internal register, the internal cache, memory Cpnt2, or storage Cpnt3. In particular embodiments, the processor Cpnt1 may include one or more internal caches for data, instructions, or addresses. The processor Cpnt1 may include one or more instruction caches and one or more data caches, for example, to hold data tables. Instructions in the instruction caches may be copies of instructions in memory Cpnt2 or storage Cpnt3, and the instruction caches may speed up retrieval of these instructions by the processor Cpnt1. Processor Cpnt1 may include any suitable number of internal registers and may include one or more arithmetic logic units (ALUs). Processor Cpnt1 may be a multi-core processor or may include one or more processors Cpnt1. Although this disclosure describes and illustrates a particular processor, this disclosure contemplates any suitable processor.

[0090] The memory Cpnt2 may include a main memory that stores instructions for the processor Cpnt1 to execute processing or to hold intermediate data during processing. For example, the computer system may load instructions or data (e.g., a data table) from the storage Cpnt3 or from other sources (e.g., another computer system) into the memory Cpnt2. The processor Cpnt1 may load instructions and data from the memory Cpnt2 into one or more internal registers or internal caches. To execute an instruction, the processor Cpnt1 may read and decode the instruction from the internal register or internal cache. During or after the execution of an instruction, the processor Cpnt1 may write one or more results (which may be intermediate or final results) to an internal register, an internal cache, the memory Cpnt2, or the storage Cpnt3. The bus Cpnt6 may include one or more memory buses (each of which may include an ADDRESS bus and a DATA bus) and may couple the processor Cpnt1 to the memory Cpnt2 and / or the storage Cpnt3. Optionally, one or more memory management units (MMUs) facilitate data transfer between the processor Cpnt1 and the memory Cpnt2. The memory Cpnt2 (which may be a high-speed volatile memory) may include random access memory (RAM), such as dynamic RAM (DRAM) or static RAM (SRAM). The storage Cpnt3 may include long-term or high-capacity storage for data or instructions. The storage Cpnt3 may be internal or external to the computer system and may include one or more of a disk drive (e.g., a hard disk drive (HDD) or a solid-state drive (SSD)), flash memory, ROM, EPROM, optical disk, magneto-optical disk, magnetic tape, a universal serial bus (USB)-accessible drive, or other type of non-volatile memory.

[0091] The I / O interface Cpnt4 may be software, hardware, or a combination of both, and may include one or more interfaces (e.g., serial or parallel communication ports) for communicating with I / O devices, which may enable communication with a human (e.g., a user). For example, the I / O devices may include a keyboard, keypad, microphone, monitor, mouse, printer, scanner, speaker, still camera, stylus, table, touch screen, trackball, video camera, other suitable I / O device, or a combination of two or more thereof.

[0092] The communication interface Cpnt5 may provide a network interface for communicating with other systems or networks. The communication interface Cpnt5 may include a Bluetooth interface or other types of packet-based communication. For example, the communication interface Cpnt5 may include a network interface controller (NIC) and / or a wireless NIC or wireless adapter for communication with a wireless network. The communication interface Cpnt5 may provide communication with a Wi-Fi network, an ad-hoc network, a personal area network (PAN), a wireless PAN (e.g., Bluetooth WPAN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), a cellular network (e.g., a Global System for Mobile Communications (GSM) network), the Internet, or a combination of two or more thereof.

[0093] Bus Cpnt6 may provide a communication link between the above-mentioned components of the computing system. For example, bus Cpnt6 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand bus, a low-pin-count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCIe) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association local (VLB) bus, or any other suitable bus, or a combination of two or more thereof.

[0094] Although this disclosure describes and illustrates a particular computer system having a particular number of particular components in a particular arrangement, this disclosure contemplates any suitable computer system having any suitable number of any suitable components in any suitable arrangement.

[0095] As used herein, a computer-readable non-transitory storage medium may include one or more semiconductor-based or other integrated circuits (ICs) (e.g., field programmable gate arrays (FPGAs) or application-specific ICs (ASICs)), hard disk drives (HDDs), hybrid hard drives (HHDs), optical disks, optical disk drives (ODDs), magneto-optical disks, magneto-optical drives, floppy diskettes, floppy disk drives (FDDs), magnetic tapes, solid-state drives (SSDs), RAM-drives, SECURE DIGITAL cards or drives, or any other suitable computer-readable non-transitory storage medium, or any suitable combination of two or more thereof, where appropriate. A computer-readable non-transitory storage medium may be volatile, non-volatile, or a combination of volatile and non-volatile, where appropriate.

[0096] While the present invention has been described in conjunction with several specific embodiments, as will be apparent to those skilled in the art in light of the foregoing description, many other alternatives, modifications, and variations will be apparent. Accordingly, the invention as described herein is intended to embrace all such alternatives, modifications, applications, and variations that may fall within the spirit and scope of the appended claims.

Claims

1. 1. A method for analyzing optical coherence tomography (OCT) data, comprising: collecting the OCT data comprising a plurality of A-scans through a medium using an OCT system; extracting a set of metrics from each individual A-scan, the set of metrics including an optical attenuation coefficient that determines how quickly the power of a coherent light beam propagating through the medium attenuates along its path, as well as one or more of RPE sub-reflectance, RPE internal reflectance, retinal thickness, and choriocapillaris flow, distance from a particular A-scan to a particular retinal structure, and layer integrity; defining a set of images based on the extracted metrics, each image defining lesion characteristic data, the lesion being geographic atrophy; defining a multi-channel image based on the set of images, wherein each metric defines a separate corresponding channel for each pixel of the multi-channel image; providing the multi-channel images to a machine learning model trained to identify one or more lesions based on the lesion characteristic data; and displaying or storing the identified lesions for future processing.

2. 10. The method of claim 1, wherein one or more of the images define pixels based on relative distances from corresponding A-scans to predetermined ocular landmarks, the pixels being based on the distance from each A-scan to the fovea.

3. 3. The method of claim 1, further comprising accessing additional imaging data of one or more additional imaging modalities different from OCT, wherein the multi-channel image comprises one or more image channels based respectively on the one or more additional imaging modalities.

4. The method of claim 3 , wherein the one or more additional images are based on one or more of a fundus image, an autofluorescence image, a fluorescence angiography image, an OCT angiography image, and a perimetry map.

5. The method of claim 1 , further comprising the step of acquiring visual field function data, wherein the multi-channel image comprises at least one image channel based on the visual field function data.

6. further comprising sorting the extracted metrics from each A-scan into corresponding metric groups; the metric group defines a lesion characteristic image; The method of claim 1 , wherein each channel of the multi-channel image is based on a corresponding metric group.

7. The OCT data includes OCT structural data and OCT angiography (OCTA) flow data; the set of metrics includes OCT-based metrics extracted from the OCT structural data and OCTA-based metrics extracted from the OCTA flow data; the set of images includes an OCT-based image based on the OCT-based metric and an OCTA-based image based on the OCTA-based metric; The method of claim 1 , wherein the multi-channel image is based on the OCT-based image and the OCTA-based image.

8. The method of claim 1 , wherein the machine learning model is embodied by a neural network.

9. The method of claim 8, wherein the neural network is a U-Net type architecture.

10. The method of claim 1 , wherein each channel in the multi-channel image is a color channel.

11. The method of claim 1 , wherein the multi-channel image is an en face image.

12. The optical attenuation coefficient is L(z)=L 0 e -μz is defined by where L(z) is the irradiance of the beam after passing through the medium over distance z, and L 0 2. The method of claim 1, wherein ∑ is the irradiance of the incident light beam and μ is the optical extinction coefficient.

Citation Information

Patent Citations

  • Attenuation-based optic neuropathy detection with three-dimensional optical coherence tomography

    JP2014147780A

  • Method for determining the depth-resolved physical and / or optical properties of a scattering medium.

    JP2014516646A

  • Identification and measurement of geographic atrophy

    JP2015136626A

  • Image analysis

    JP2018515297A

  • Ophthalmologic image processing device and ophthalmologic image processing program

    JP2019177032A