Segmentation and Classification of Geographic Atrophy Patterns in Patients with Age-Related Macular Degeneration in Wide-Field Autofluorescence Images

A two-stage segmentation process with supervised classifiers and dynamic contour algorithms, enhanced by deep learning, automates the identification and classification of GA phenotypes in fundus autofluorescence images, addressing the limitations of semi-automatic methods and improving GA lesion analysis efficiency.

JP7712204B2Active Publication Date: 2025-07-23CARL ZEISS MEDITEC INC +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2021538343
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-02-08
Filing Date
2020-02-06
Publication Date
2025-07-23
Estimated Expiration
2040-02-06

AI Technical Summary

Technical Problem

Current segmentation algorithms for geographic atrophy (GA) lesions in fundus autofluorescence images are semi-automatic and require manual input, and there is a need for a fully automatic method to quantify GA lesions and classify their phenotypic patterns, especially in wide-field images where landmarks are not well defined.

Method used

A two-stage segmentation process using a supervised classifier and a dynamic contour algorithm, combined with deep learning neural networks, to identify and classify GA regions as 'diffuse' or 'striated' phenotypes based on contour non-uniformity and intensity uniformity measures, respectively.

Benefits of technology

Enables fully automated and accurate segmentation and classification of GA phenotypes, reducing manual intervention and improving the efficiency of GA lesion analysis in wide-field images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007712204000001
    Figure 0007712204000001
  • Figure 0007712204000002
    Figure 0007712204000002
  • Figure 0007712204000003
    Figure 0007712204000003
Patent Text Reader

Abstract

An automated segmentation and classification system / method for identifying geographic atrophy (GA) phenotypes in fundus autofluorescence images. The hybrid process combines a supervised pixel classifier with an active contour algorithm. A trained machine learning model (e.g., SVM or U-Net) provides the initial GA segmentation / classification, followed by the application of the Chan-Vese active contour algorithm. Junction zones of the GA segmentation area are analyzed for geometric and light intensity regularities. GA phenotype identification is at least partially determined from these parameters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to the field of ophthalmic autofluorescence images. More particularly, it relates to the classification of geographic atrophy regions in fundus autofluorescence images.

Background Art

[0002] Age-related macular degeneration (AMD) is the leading cause of blindness in the elderly population in developing countries. Geographic atrophy (GA) is an advanced form of AMD, characterized by the loss of photoreceptor cells, retinal pigment epithelium (RPE), and choroidal capillaries. The number of patients with GA is estimated to be 5 million worldwide. GA can cause irreversible loss of visual function and accounts for approximately 20% of severe visual impairment from AMD. There is currently no approved treatment to counter the progression of GA. However, in recent years, the understanding of the etiology of GA has advanced, and several promising treatment clinical trials are underway. Nevertheless, early identification of GA progression is extremely important in delaying its effects.

[0003] Differences in the expression patterns of abnormal fundus autofluorescence (FAF) images have been found to be useful in identifying the progression of GA. Therefore, the identification of GA lesions and their phenotypes can be an important factor in the identification of the disease progression and clinical diagnosis of AMD. Although it is also effective to visually inspect FAF images of GA lesions and their phenotypes using medical staff, it takes time. Several segmentation algorithms have been developed to assist in the evaluation of GA lesions, but many of these algorithms are semi-automatic and require manual input for the segmentation of GA lesions.

Summary of the Invention

Problems to be Solved by the Invention

[0004] An object of the present invention is to provide a fully automatic method for quantifying GA lesions and their expression patterns in medical images. Another object of the present invention is to supplement the GA lesion segmentation algorithm with the ability to identify and classify different phenotypic patterns.

[0005] Another object of the present invention is to provide a framework for automatic segmentation and identification of GA expression patterns in wide-field FAF images.

Means for Solving the Problems

[0006] The above object is achieved in a system / method for classifying (e.g., identifying) geographic atrophy (GA) in the eye. For example, the GA region (e.g., GA segmentation in a fundus or en face image) may be identified as a "diffuse" phenotype or a "striated" phenotype. Both of these phenotypes have empirically been identified as indicating geographic atrophy with a rapid progression rate. The system may include an ophthalmic diagnostic device such as a fundus imager or an optical coherence tomography (OCT) device used to generate / capture ophthalmic images. Alternatively, the ophthalmic diagnostic device may embody a computer system that accesses existing ophthalmic images from a data store using a computer network such as the Internet. The ophthalmic diagnostic device is used to acquire fundus images. Preferably, the image is a fundus autofluorescence (FAF) image because GA lesions are typically more recognizable in such images.

[0007] The acquired image is then provided to an automatic GA identification process, and the GA regions in the image are identified (e.g., segmented). The GA segmentation process of the present application may be fully automated and based on a deep learning neural network. The GA segmentation may further be based on a two-stage segmentation process, for example, a hybrid process combining a supervised classifier (e.g., pixel / image segment segmentation) and a dynamic contour algorithm. The supervised classifier is preferably a machine learning (ML) model and may be implemented as a support vector machine (SVM) or a deep learning neural network, preferably a U-Net type convolutional neural network. This first-stage classifier / segmentation ML model identifies the initial GA regions (lesions) in the image, and the results are fed into a dynamic contour algorithm for the second-stage segmentation. The dynamic contour algorithm may be implemented as a modified Chase-Vese segmentation algorithm, and the initial GA regions identified in the first stage are used as the starting points (e.g., initial contours) in the Chase-Vese segmentation. As a result, robust GA segmentation of the acquired image becomes possible.

[0008] The identified GA regions are then provided for analysis to identify those specific phenotypes. For example, the system / method may identify a "diffuse" phenotype using a contour non-uniformity measure of the GA region and identify a "striped" phenotype using an intensity uniformity measure. Both of these two phenotypes indicate a high progression rate GA region. This analysis may be performed in a two-stage process, where the first of the two measures is calculated, and if the first measure is higher than a first threshold, the GA region may be classified as having a high progression rate, and there is no need to calculate the second of the two measurements. However, if the first measure is below the first threshold, the second measure is calculated and compared to a second threshold to determine whether it indicates a high progression rate GA. For example, if the contour non-uniformity measure is higher than a first predetermined threshold (e.g., the non-uniformity of the peripheral contour of the GA region is higher than the predetermined threshold), the GA region may be analyzed as having a "diffuse" phenotype, and if the intensity uniformity measure is higher than a second threshold, the GA region may be classified as having a "striped" phenotype.

[0009] Other objects and attainments of the present invention will become apparent and be understood by reference to the following description and the claims, taken in conjunction with the accompanying drawings, which will provide a more complete understanding of the present invention.

[0010] The embodiments disclosed herein are merely examples and the scope of the present disclosure is not limited thereto. The features of any embodiment described in one claim category, e.g., a system, can also be claimed in other claim categories, e.g., a method. The dependencies or back-references in the appended claims are selected only for formal reasons. However, any subject matter obtained from careful back-references to the previous claims can also be claimed, thereby disclosing any combination of claims and their features, and can be claimed regardless of the dependencies selected in the appended claims.

[0011] In the figures, like reference symbols / characters refer to like components.

Brief Description of the Drawings

[0012]

Figure 1

Figure 2

Figures 3A - 3B

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

[0013] Early diagnosis is extremely important for the successful treatment of various eye diseases. Optical imaging is a preferred method for non-invasive examination of the retina. Age-related macular degeneration (AMD) is known to be a major cause of blindness, but diagnosis is often made only after the damage itself has occurred. Therefore, the purpose of advanced ophthalmic imaging devices is to provide diagnostic tools that can detect and monitor pathological changes at the pre-clinical stage of the disease.

[0014] Several ophthalmic imaging systems are known in the art, such as fundus imaging systems and optical coherence tomography systems, any of which may be used in the present invention. Examples of ophthalmic imaging modalities are provided below, see for example FIGS. 13 and 14. Any of these devices may be used to provide a fundus image, where the fundus is the inner surface of the eye opposite the lens (intraocular lens) and may include the retina, optic disc, macula, fovea, and posterior pole.

[0015] An ophthalmic imaging system can generate full-color images. Other imaging techniques, such as fluorescein angiography or indocyanine green angiography (ICG), may be used to capture high-contrast images of specific ophthalmic features such as blood vessels and lesions. High-contrast images are obtained by injecting a fluorescent dye into the subject's bloodstream and then collecting an image using a specific light frequency (e.g., color) selected to excite the fluorescent dye for image capture.

[0016] Alternatively, high-contrast images may be obtained without using a fluorescent dye. For example, in a technique known as autofluorescence, individual light sources that provide specific frequencies of light (e.g., colored LEDs or lasers) can be used to excite different fluorophores that occur naturally within the eye. The resulting fluorescence can be detected by removing the excitation wavelength with a filter. This can make it possible to recognize some features / tissues of the eye more easily than at the level possible with true-color images. For example, fundus autofluorescence (FAF) imaging can be performed by green or blue excitation that stimulates the natural fluorescence of lipofuscin, and a monochrome image is generated. For illustrative purposes, FIG. 1 provides three grayscale images 11A, 13A, and 15A of a true-color image and, adjacent thereto, three corresponding autofluorescence images 11B, 13B, and 15B of the same region of the eye. As is apparent, some tissue features that are not very sharply defined within the grayscale images 11A, 13A, and 15A (e.g., within the central area of each image) are readily distinguishable in the corresponding autofluorescence images 11B, 13B, and 15B.

[0017] AMD generally begins as dry (atrophic) AMD and progresses to wet (neovascular) AMD. Dry AMD is characterized by drusen (e.g., small white or yellowish deposits) that form on the retina (e.g., under the macula), causing the retina to deteriorate or degenerate over time. Early dry AMD may be treated with nutritional therapy or supplements, and wet AMD may be treated with injections (e.g., injections of Lucentis, Avastin, and Eylea), but for more advanced dry AMD, while clinical trials are underway, there are no approved treatments at present. Geographic atrophy (GA) is a late form of dry AMD, characterized by patches of deteriorated or dead cells within the retina. GA may be defined as any sharply demarcated round or oval area of depigmentation, i.e., a distinct loss of the retinal pigment epithelium (RPE), where choroidal defects are more visible than in the surrounding area and which is of a minute size, e.g., 175 μm in diameter. Historically, color fundus images have been used for imaging and discrimination of GA, but color fundus images cannot visualize the features of lesions associated with the progression of GA. In fundus autofluorescence (FAF) imaging, the atrophic areas appear dark due to the loss of RPE cells containing the autofluorescent dye lipofuscin. This results in a high contrast between the atrophic and non-atrophic areas, making it easier to demarcate the area of GA than with color fundus images. The use of autofluorescence has been found to be a more effective imaging modality for the assessment of GA lesions and for monitoring changes in lesion structure. In particular, FAF with an excitation wavelength of 488 nm is the current industry-preferred technique for morphological assessment of GA lesions. For illustrative purposes, FIG. 2 provides an exemplary FAF image of a GA lesion.

[0018] The FAF image of GA may be characterized by an abnormal pattern of excessive autofluorescence surrounding the atrophy area. This may lead to classification or identification of a specific phenotype of GA. The classification into these different patterns of characteristic excessive autofluorescence was first presented by Bindewald et al. in "Classification of Fundus Autofluorescence Patterns in Early Age-Related Macular Disease", Invest Ophthalmol Vis Sci, 2005, Vol. 46, pp. 3309-14.Subsequently, multiple research results have been published, demonstrating the impact of different phenotypic patterns on the progression of the disease and their ability to serve as a prognostic determinant, as described in Holz F.G. et al., "Progression of Geographic Atrophy and Impact of Fundus Autofluorescence Patterns in Age-Related Macular Degeneration," Am J Ophthalmol, 2007, Vol. 143, pp. 463-472; Fleckenstein M. et al., "The ‘Diffuse-Trickling’ Fundus Autofluorescence Phenotype in Geographic Atrophy," Invest Ophthalmol Vis Sci., 2014, Vol. 55, pp. 2911-20; and Jeong Y.J. et al., "Predictors for the Progression of Geographic Atrophy in Patients with Age-Related Macular Degeneration: Fundus Autofluorescence Study with Modified Fundus Camera," Eye, 2014 online, 28(2), pp. 209-218, DOI 10.1038 / eye.2013.275. The entire contents of the above references are incorporated herein by reference.

[0019] Of particular interest are the "diffuse" (or "diffuse-trickling") GA phenotypes and the "banded" GA phenotypes, both of which have been empirically identified as exhibiting a higher rate of worsening of GA. Therefore, in addition to identifying the GA regions, it is equally important to classify the appropriate phenotypes of the identified GA regions. The "diffuse" GA regions may be characterized by the degree of inhomogeneity of their surrounding contours (e.g., a contour inhomogeneity measure higher than a predetermined inhomogeneity threshold). As a result, the GA region in FIG. 2 can be identified as not being of the "diffuse" phenotype, as its contour is relatively uniform. In contrast, FIG. 3A provides some FAF image examples of GA lesions of the "diffuse" phenotype, all of which exhibit higher contour inhomogeneity. Since the concentration of excessive autofluorescence surrounding the atrophy region may be high in the GA region, the "banded" GA regions may be characterized by the amount of variation in the intensity of excessive autofluorescence along their periphery. More specifically, if the variation in the intensity of the contour excessive autofluorescence in the GA region is relatively uniform (e.g., an intensity uniformity measure higher than a predetermined intensity threshold), the GA region may be characterized as the "banded" phenotype. For example, the GA region in FIG. 2 can be identified as not being of the "banded" phenotype, as it lacks uniformity of excessive autofluorescence along its contour. In comparison, FIG. 3B provides an example of FAF of a GA lesion of the "banded" phenotype in which uniform excessive autofluorescence is seen along its contour.

[0020] Automated quantification of GA lesions and their phenotypic patterns may be highly useful for identifying disease progression and for clinical diagnosis of AMD. As described above, currently there are several segmentation algorithms for GA lesion evaluation, but none of them are capable of identifying and classifying different phenotypic patterns. That is, to date, none of the available segmentation methods for GA lesions are able to automatically classify different GA phenotypic patterns. Most of the segmentation methods reported in the literature are semi-automated and require manual input for segmentation of GA lesions. Furthermore, all of the previously reported segmentation methods were developed for use with standard field of view (FOV) images (e.g., FOV of 45° - 60°), and the location of GA lesions was generally limited to a predetermined area of the standard FOV image. However, due to the increasing interest in recent wide-angle images (e.g., FOV of 60° - 120° or more), the need for an automated GA evaluation method suitable for use with wide-angle images has been emphasized. The use of wide-angle images complicates the use of typical GA segmentation algorithms, as the location of GA lesions cannot be limited to a predetermined area of the wide-angle image and some (physical) landmarks in the diseased eye may not be well defined.

[0021] Figure 4 provides a schematic framework for the automated segmentation and identification of GA phenotypic patterns. First, a measure of image quality (IQ) is specified (block B1). This may include applying an image quality algorithm to the FAF image to reject images of ungradeable quality (IQ measure below a predetermined value). If an image is rejected, the process may jump directly to B9. If the IQ measure is sufficient to proceed with grading, the process continues to block B3.

[0022] Block B3 provides optional preprocessing for the FAF image. This may include detection of the optic nerve head and / or blood vessels, selection of the region of interest (ROI), and histogram correction. The automatic ROI selection may be based on the detection of the optic nerve head, limiting the amount of pixels that require processing (e.g., classification), thereby shortening the execution time of the segmentation process.

[0023] Block B5 provides GA lesion segmentation. This may combine pixel-by-pixel classification with an improved Chan-Vese active contour. For example, the classification may be provided by a machine learning (ML) model (e.g., support vector machine (SVM) and / or deep learning neural network), which provides an initial GA segmentation. Then, the GA segmentation is supplied as a starting point for the Chan-Vese active contour algorithm that further refines the GA segmentation. This hybrid process utilizes a novel combination of a supervised pixel classifier and an active contour.

[0024] The GA segmentation of block B5 is then provided to the phenotypic classification block B7, which identifies and classifies the junction zones (e.g., along the peripheral junctions of the GA segmentation) in the vicinity of the GA segment area. For example, a set of random points that are equidistant from each other and distributed along the periphery of the GA segmentation may be selected. The distance from each selected point to the centroid of the GA segmentation may be calculated. These distances may then be used to specify the peripheral contour smoothness of the GA segmentation (e.g., the smoothness of the contour traced by GA). The peaks and valleys of intensity may be calculated along the direction perpendicular to the peripheral contour of the GA segmentation and outward, for example, using the gradient differential coefficient processed by a Hessian filter and / or a directional Gaussian filter. This may provide a measure of the light intensity regularity along the periphery of the GA segmentation. Both of these parameters may be used to classify the phenotype of the junction zone. For example, the measure of the contour smoothness of the GA segmentation may be used to identify the "diffuse" phenotype, and the measure of the intensity regularity of the GA segmentation may be used to identify the "striped" phenotype.

[0025] Alternatively, a (e.g., deep learning) neural network may be implemented as block B7. In this case, the neural network may be trained to receive the GA segmentation of block B5 and classify the specific phenotype (e.g., diffuse or striped) of each GA segmentation. Further alternatively, the neural network may be trained to perform the functions of the GA lesion segmentation block B5 and the phenotypic classification block B7. For example, a set of images showing the GA regions drawn by an expert and their respective phenotypes identified by the expert may be used as the training output set for the neural network, and the initially undrawn images may be used as the training input set for the neural network. The neural network may thereby be trained not only to segment (or draw) the GA lesion regions but also to identify their respective phenotypes (e.g., "diffuse" or "striped").

[0026] In addition, it has been found that retinal blood vessels may cause some fundus regions to be misidentified as GA segmentations. This may be avoided by excluding retinal blood vessels (e.g., the retinal blood vessels identified in preprocessing block B3) from the FAF image before applying the GA segmentation, for example, before block B5. In this case, the retinal blood vessels may optionally be excluded from the training images used to train the learning model (SVM or neural network).

[0027] The last block B9 generates a report summarizing the results of the current framework. For example, the report may specify whether the image quality was too low to be graded, indicate whether the identified GA lesion is of the "high" worsening rate type, indicate the phenotype of the GA lesion, and / or propose a follow-up date. Both the "diffuse" and "striate" phenotypes have been empirically identified as indicating high-worsening rate GA, but the "diffuse" phenotype may have a faster worsening rate than the "striate" phenotype. The proposed follow-up date may be based on the identification of the type of GA lesion. For example, GA lesions with a higher worsening rate may have a legitimate reason for an earlier follow-up date than GA lesions of a lower or "low" (or not high) worsening rate type.

[0028] FIG. 5 illustrates a more detailed process for the automatic segmentation and identification of GA phenotype patterns according to the present invention. For purposes of illustration, the present invention is described as being applied to wide-angle FAF images, but may be similarly applied to other types of ophthalmic imaging modalities, such as OCT / OCTA, that can generate images providing visualization of, for example, GA. For example, this may be applied to en face OCT / OCTA images. The imaging system and / or method of the present application may begin by capturing, or otherwise obtaining (e.g., by accessing its data store), a fundus autofluorescence image (step S1). Next, an assessment (e.g., measurement) of the image quality of the obtained fundus autofluorescence image is performed using any suitable image quality (IQ) algorithm known in the art. For example, the measure of image quality may be based on one or a combination of quantifiable elements such as sharpness, noise, dynamic range, contrast, vignetting, etc. If the measure of image quality is lower than a predetermined image quality threshold (step S3 is yes), the FAF image is identified as not suitable for subsequent analysis, and the process proceeds to step S29, where a report is generated stating that GA segmentation or quantification cannot be performed on the current FAF image. In essence, the image quality IQ algorithm is applied to reject images of non-gradable quality. If the image quality is above the minimum threshold (step S3 = no), the process may proceed to an optional preprocessing step S5, or alternatively may proceed directly to the GA image segmentation step S7.

[0029] Processing step S5 may include a plurality of preprocessing sub-steps. For example, it may include optic disc and / or blood vessel detection (e.g., a finding mask), and in addition, it may be specified whether the image is of the patient's left or right eye. As described above, it may be difficult to identify physical landmarks in the diseased eye, but the optic disc may be identifiable by the concentration of the vascular structure emerging therefrom. For example, a machine learning model (e.g., a deep learning neural network as described below) may be trained to identify the optic disc in a wide-angle image. Alternatively, other types of machine learning models (e.g., support vector machines) may be trained to identify the optic disc, for example, by correlating the position of the optic disc with the position where the concentration of the vascular structure is located. In either case, additional information other than the image, for example, information provided by an electronic medical record (EMR) or by the "Digital Imaging and Communications in Medicine" (DICOM) standard, may be used in the training.

[0030] As described above, it has been found that retinal blood vessels may cause false positives in GA segmentation. Therefore, the preprocessing may further identify and exclude retinal blood vessels before applying GA segmentation to reduce the number of false positives, for example, to reduce the identification of incorrect GA regions. As described above, since GA segmentation may be based on a machine learning model, retinal blood vessels may be excluded from its training image set. For example, a neural network may use a training output set of fundus images (e.g., FAF images) with manually delineated GA regions and manually identified phenotypes, and may also be trained using a training input set of the same fundus images that do not have manually delineated GA regions and do not have identified phenotypes. In a certain operation phase (or test phase), the neural network will accept input test images from which retinal blood vessels have been excluded, and then the neural network may be trained with training images from which the retinal blood vessels have also been excluded. That is, before training the neural network, retinal blood vessels may be excluded from the fundus images in the training input set and the training output set.

[0031] The preprocessing step S5 may also include the selection of a region of interest (ROI) to limit the processing (including GA segmentation) to the identified ROI in the FAF image. Preferably, the identified ROI will include the retinal macula. ROI selection may be automated based on landmark detection (optic nerve head and retinal blood vessels), image size, image entropy, and fundus scan information extracted from DICOM, etc. For example, depending on whether the image is of the left eye or the right eye, the retinal macula is to the right or left of the optic nerve head in the image. As will be understood, ROI selection limits the amount of pixels that need to be classified (e.g., for GA segmentation), thereby shortening the execution time of the process / method of the present application.

[0032] Processing step S5 may further include illumination correction and contrast adjustment. For example, non-uniform image illumination may be corrected using background removal, and the image contrast may be adjusted using histogram equalization.

[0033] Next, the processed image is provided to a two-stage lesion segmentation (e.g., block B5 in FIG. 4), which includes GA segmentation and dynamic contour analysis. The first stage of block B5 is the GA segmentation / classifier stage (S7), which is preferably based on machine learning and identifies one or more GA regions. The second stage of block B5 is the dynamic contour algorithm applied to the identified GA regions. The GA segmentation stage S7 may provide GA classification for each sub-image sector, where each sub-image sector may be one pixel (e.g., per pixel), or each sub-image sector may be a set of multiple pixels (e.g., a window). In this way, each sub-image sector may be individually classified as a GA sub-image or a non-GA sub-image. The dynamic contour analysis stage S9 may be based on the known Chan-Vese algorithm. In this case, the Chan-Vese algorithm is modified to change the energy and movement direction of contour growth. For example, the contour movement is determined using the image characteristics for each region. Essentially, the segmentation block B5 for GA lesions utilizes a hybrid process that combines (e.g., per pixel) classification and the Chan-Vese dynamic contour. Optionally, the hybrid algorithm proposed here utilizes a novel combination of a supervised (e.g., pixel) classifier (S7) and a geometric dynamic contour model (S9). As will be understood by those skilled in the art, a geometric dynamic contour model typically starts from a contour (e.g., a starting point) that defines an initial segmentation within the image plane and then develops the contour according to some evolution equation until it stops at the boundary of the foreground region. In this case, the geometric dynamic contour model is based on the modified Chan-Vese algorithm.

[0034] The GA segmentation / classifier stage (S7) may preferably be implemented based on machine learning, for example, by using a support vector machine or by a (deep learning) neural network. In this specification, each implementation will be described separately.

[0035] Generally, a support vector machine, SVM, is a machine learning linear model for classification and regression problems and may be used to solve linear and non-linear problems. The idea of SVM is to create a line or hyperplane that separates data into classes. More formally, an SVM defines one or more hyperplanes in a multi-dimensional space, and the hyperplanes are used for classification, regression, outlier detection, etc. Essentially, an SVM model represents labeled training examples as points in a multi-dimensional space, and the labeled training examples of different categories are mapped such that they are separated by a hyperplane that may be considered a decision boundary for separating different categories. When a new test input sample is supplied to the SVM model, this test input is mapped into the same space, and a prediction regarding which category it belongs to is made based on which side of the decision boundary (hyperplane) the test input lies.

[0036] In a preferred embodiment, the SVM is used for image segmentation. Image segmentation aims to divide an image into different sub-images with different characteristics and extract the objects of interest. More specifically in this specification, the SVM is trained to segment the GA regions in the FAF image (for example, trained to clearly define the contours of the GA regions in the FAF image). Various SVM architectures for image segmentation are known in the art, and the specific SVM architecture used for this task is not critical to the present invention. For example, a least squares SVM may be used for image segmentation based on per-pixel (or per-sub-image sector) classification. Both pixel-level features (such as color, intensity, etc.) and texture features may be used as inputs to the SVM. Optionally, an ensemble of SVMs, each providing a specialized classification, may be associated to obtain good results.

[0037] Therefore, the initial contour selection may be performed using an SVM classifier (e.g., by segmentation based on an SVM model). Preferably, Haralick texture features, average intensity, and variance parameters obtained from a gray-level co-occurrence matrix (e.g., an 11×11 window moving within the region of interest) are used for training the SVM classifier. As described above, the retinal blood vessels may optionally be excluded from the training images. In any case, the feature extraction is limited to a specific ROI (e.g., the ROI selected in step S5), and as a result, better time performance is obtained than when GA segmentation / classification is applied to the entire image. In this way, the SVM provides an initial contour selection that is the target of the dynamic contour algorithm in step S9 (e.g., provides an initial GA segmentation as a starting point) for better performance. The development time of the dynamic contour algorithm depends greatly on the initial contour selection, which is described, for example, in Chen, Q, et al., "Semi-automatic geographic atrophy segmentation for SD-OCT images", Biomedical Optics Express, Vol. 4(12), pp. 2729-2750, which is hereby incorporated by reference in its entirety.

[0038] Tests and comparisons were made regarding the performance of the SVM classifier. An expert grader manually delineated the GA regions in FAF images not used for training, and the results were compared with the segmentation results by the SVM model. FIG. 6 shows four GA regions delineated by a human expert, and FIG. 7 shows the corresponding GA delineation (e.g., around the GA segmentation) provided by the SVM of the present application. As shown in the figures, the SVM model is in good agreement with the manual GA segmentation provided by the expert grader.

[0039] The GA segmentation / classifier stage (S7) may also be implemented by using a neural network (NN) machine learning (LM) model. Various examples of neural networks will be described below with reference to FIGS. 16-19, any of which, or combinations thereof, may be used in the present invention. An exemplary implementation of a GA segmentation / classifier using a deep learning neural network was configured based on the U-Net architecture (see FIG. 19). In this case, the NN was trained using manually segmented images (e.g., images in which the GA segments were segmented by a human expert) as training outputs and the corresponding unsegmented images as training inputs.

[0040] In one exemplary embodiment, 79 FAF green images were obtained from 62 GA patients using a CLARUS™ 500 fundus camera (Zeiss, Dublin, California). These 79 FAF images were divided into 55 FAF images for training and 24 FAF images for testing. Optionally, retinal blood vessels may be excluded from the training and / or test images. A data augmentation method was used to increase the scale of the training data, and 880 (image) patches of size 128×128 pixels were generated for training. In this U-Net, the downsampling path was composed of four convolutional neural network (CNN) blocks, each CNN block being composed of two CNN layers followed by one max pooling layer. The bottleneck (e.g., the block between the downsampling path and the upsampling path) was composed of two CNN layers and an optional dropout of 0.5 to avoid the problem of overfitting. The upsampling path of the U-net is typically symmetric to the downsampling path and was here composed of four CNN blocks followed by a bottleneck. Each CNN block of the upsampling path was composed of a transposed convolutional layer, a concatenation layer, and two subsequent CNN layers. The last CNN block (e.g., the fourth CNN block) of this upsampling path provided a segmentation output, which may be fed to an optional classification layer for labeling. This last classifier formed the last layer. In this embodiment, a custom dice coefficient loss function was used to train the machine learning model, although other loss functions, such as the cross entropy loss function, may also be used. The segmentation performance of the DL ML model may be fine-tuned by using post-training optimization. For example, a method called "Icing on the Cake" may be used, which retrains only the last classifier (i.e., the last layer) after the end of the initial (usual) training.

[0041] The DL method was compared with the manually segmented GA regions. Figure 8 shows five GA regions manually drawn by an expert grader, and Figure 9 shows the corresponding GA drawing provided by the DL machine model of the present application. For evaluation, the fractional area difference, overlap ratio, and Pearson's correlation coefficient between the measured areas were identified. The fractional area difference between the GA regions generated by the DL machine model and the manual segmentation was 4.40% ± 3.88%. The overlap ratio between the manual and DL automatic segmentation was 92.76 ± 5.62, and the correlation coefficient of the GA areas generated by the DL algorithm and the manual grading was 0.995 (p-value < 0.001). Thus, the quantitative and qualitative evaluations demonstrate that the proposed DL model for segmentation of GAs in FAF images shows a very strong agreement with the manual grading by experts.

[0042] Therefore, as is apparent from FIGS. 6 to 9, both the SVM-based and DL-based classifications provide good initial GA drawings or segmentations. The GA segmentation obtained from step S7 may be supplied as an initial contour selection for the active contour algorithm (step S9). This embodiment uses an improved Chan-Vese (C-V) active contour segmentation algorithm. Generally, active contour algorithms have been used for GA segmentation in optical coherence tomography (OCT) images, as described in Niu, S., et al., "Automated geographic atrophy segmentation for SD-OCT images using region-based C-V model via local similarity factor," Biomedical Optics Express, February 1, 2016, Vol. 7(2), pp. 581-600, which is hereby incorporated by reference in its entirety. However, in this case, the Chan-Vese active contour segmentation is used as the second phase of a two-stage segmentation process. The initial contour provided by the SVM or DL learning model classifier is close to the GA boundary annotated by an expert. This reduces the execution time for the Chan-Vese active contour segmentation and improves its performance.

[0043] The final GA segmentation result provided by the two-stage segmentation block B5 may optionally be provided to additional morphological operations (step S11) to further refine the segmentation. Morphological operations (e.g., erosion, dilation, opening, and closing) are applied to the output of the Chan-Vese active contour to refine the contour boundary and remove small isolated regions.

[0044] Next, the size of the segmented GA is determined. Although GA segmentation is obtained from a 2D image, since the fundus is curved, the 2D GA segmentation may be mapped into 3D space (step S13), and the measurement may be performed in 3D space to address distortion and obtain a more accurate area measurement. Any suitable mapping method based on known imaging geometry from 2D pixels to 3D coordinates may be used. For example, a well-known method in the art for mapping pixels in a 2D image plane to points on a sphere (e.g., 3D space) is stereo projection. Another example of 3D reconstruction from a 2D fundus image is provided in Cheriyan, J. et al., "3D Reconstruction of Human Retina from Fundus Image - A Survey", International Journal of Modern Engineering Research, Vol. 2, No. 5, September - October 2012, pp. 3089 - 3092, which is hereby incorporated by reference in its entirety.

[0045] In determining the GA size measurement value in 3D space, additional information provided by an ophthalmic imaging system may be used if available. For example, some ophthalmic imaging systems store the 3D position within the array along with the 2D coordinates of each pixel in the 2D image. This method may be included in a wide - angle ophthalmic photographic image module that conforms to the Digital Imaging and Communications in Medicine (DICOM) standard in healthcare. Alternatively, or in combination, some systems may store a model that can be executed for any set of 2D positions to generate the associated 3D positions.

[0046] If the determined area in 3D space is smaller than a predetermined threshold (step S15 = yes), this GA segmentation is determined to be too small to be an actual GA lesion, and the process proceeds to step S29 to generate a report. If the determined area is a predetermined threshold, e.g., 96Kμm2 If it is smaller (step S15 = no), the GA segmentation is recognized as an actual GA lesion, and the process attempts to identify the phenotype of the identified GA segmentation.

[0047] The first step in phenotype identification may be to analyze the contour smoothness of the GA segmentation (step S17). This may include several sub-steps. Identifying and classifying the junction zones close to the GA segmented area (e.g., along the edges of the GA segment area) may begin by first identifying the centroid of the GA segment area. This centroid calculation may be applied to the 2D image representation of the GA segment area. FIG. 10 illustrates a diffuse GA area 21, and its centroid is indicated by a circle 23. The next step may be to identify a set of randomly spaced points equally spaced along the perimeter of the GA segmentation. FIG. 11 illustrates a striped phenotype GA area 25, having a centroid 27 and a set of 13 randomly spaced points 29 equally spaced along the perimeter of the GA segmentation (each identified as the center of a white crosshair). It should be understood that 13 is an exemplary value, and any number of points sufficient to surround the perimeter of the GA area may be used. Next, the distance from each point 29 to the centroid 27 is calculated (e.g., the straight-line distance 31). The deviation of the distance 31 is used to distinguish a "diffuse" phenotype from others such as a "striped" phenotype.

[0048] In step S19, when the distance deviation is greater than the distance deviation threshold (step S19 = yes), this GA region is identified as a "diffused phenotype" (step S21), and the process proceeds to step S29, where a report is generated. This distance deviation, or measure of contour non-uniformity, may be specified as the average distance deviation for the current GA segmentation, and the distance deviation threshold may be defined as 20% of the standardized average deviation of (e.g., diffused) GA segments. The standardized average deviation may be specified from a library of GA segments, e.g., a (standardized) library of FAF images of GA lesions. However, when the distance deviation is less than or equal to the distance deviation threshold (step S19 = no), the process proceeds to step S23.

[0049] In step S23, along a direction perpendicular to the GA contour, peaks and valleys of intensity are identified for a set (at least a part) of the random points 29. This may be done using a Hessian filter processed gradient derivative and / or a directional Gaussian filter. If peaks and valleys (light intensity regularity - decreasing fluorescence (hypoflourences)) greater than the current intensity deviation threshold are present for a percentage (e.g., 60%) higher than a predetermined percentage of the selected points 29, (step S25 = yes), the image (e.g., GA segmentation) is classified as a "striped" phenotype (step 27), and the process proceeds to step S29, where a report is generated. The peaks and valleys of intensity may define a measure of intensity variability, and the intensity deviation threshold may be defined as 33% of the standardized average intensity deviation of (e.g., striped) GA segments. This standardized average intensity variation may be specified from a library of GA segments, e.g., a (standardized) library of FAF images of GA lesions. If the light intensity regularity - decreasing fluorescence is less than or equal to the current intensity deviation threshold (step 25 = no), no phenotype can be specified, and the process proceeds to step S29.

[0050] As described above, the neural network may be trained to provide GA segmentation as described above with respect to step S7. However, a neural network such as the U-Net of FIG. 19 may also be trained to provide a phenotypic classification. For example, by providing an expert labeling of each particular phenotype of the GA lesion segmentation drawn by an expert in a training image set, the neural network may be trained to provide a phenotypic classification, along with or in addition to, GA lesion segmentation identification. In this case, the phenotype identification / classification steps associated with steps S17 to S25 may be omitted and provided by the neural network. Alternatively, the phenotypic classification provided by the neural network may be combined with the phenotype identification results of steps S17 to S25, for example, by weighted averaging.

[0051] In step S29, a report summarizing the results of this process is generated. For images with ungradeable IQ (step S3 = yes), only the image quality is reported, and for other images, the generated report includes one or more of the following: a. The measured GA segmentation area in this visit.

[0052] b. The current GA junction zone phenotype (if available). c. The risk of worsening (e.g., the risk of worsening reported as high for "striate" and "diffuse" phenotypes).

[0053] d. A proposal for the follow-up visit time. Figure 12 illustrates another exemplary method for the automatic classification of geographic atrophy (GA) within the eye. In this example, rather than checking each of the specific GA phenotypes within a phenotypic library (e.g., the "diffuse" phenotype and the "striate" phenotype), the method stops classifying the current GA region as soon as the current GA region is identified as any phenotype associated with a high rate of worsening (e.g., an empirical association). The current GA region is then simply classified as "high-worsening rate" GA, and its specific GA phenotype is not specified. Optionally, at will, by using user input, for example, via a graphical user interface (GUI), the process may proceed to classify the current GA segmentation as one or more specific phenotypes. In this case, the method may further identify the "high-worsening rate" GA as either a "diffuse" phenotype or a "striate" phenotype. Optionally, the method may order an examination of GA phenotypes such that the phenotype is more common in a particular population (e.g., the population to which a patient may belong), or order an examination based on which the phenotype is identified (e.g., empirically) as being associated with a higher rate of worsening GA than others. For example, the "diffuse" phenotype GA may plausibly be considered to have a faster rate of worsening than the "striate" phenotype GA based on empirical observations, and thus the method may first check whether the current GA segmentation is the "diffuse" phenotype, and only if the current GA segmentation is determined not to be the "diffuse" phenotype, second check whether it is the "striate" phenotype.

[0054] The method may begin in method step M1 by using an ophthalmic diagnostic device to acquire an image of the fundus. The image (e.g., an autofluorescence image or an en face image) may be generated by a fundus imager or an OCT. The ophthalmic diagnostic device may be a device that generates the image, or alternatively, a computing device that accesses an image from a data store of such images on a computer network (e.g., the Internet).

[0055] In method step M3, the acquired image is provided to an automatic GA identification process that identifies the GA region within the image. The automatic GA identification process may include a GA segmentation process, as is known in the art. However, a preferred segmentation process is a two-step segmentation that combines GA classification (e.g., pixel-wise) with dynamic contour segmentation. The first of this two-step process may be a trained machine model such as an SVM or (e.g., deep learning) neural network that segments / classifies / identifies the GA region within the image. For example, a neural network based on the U-Net architecture may be used, and its training set may include a training output set of fluorescence images defined by an expert and a corresponding training input set of fluorescence images without defined limits. Regardless of the type of GA segmentation / classification used at this initial stage, the identified GA segmentation may be supplied as a starting point to a dynamic contour algorithm (e.g., Chan-Vese segmentation) to further refine the GA segmentation. Optionally, the result may be provided to morphological operations (e.g., shrinking, dilation, opening, and closing) to further clean the segmentation before proceeding to the next method step M5.

[0056] Optionally, before proceeding to step M5, the final GA segmentation from step M3 may be mapped into a three-dimensional space representing the shape of the fundus, and the area of the mapped GA is determined. If the area is smaller than a predetermined threshold, the GA segmentation may be reclassified as non-GA and excluded from subsequent processing. If the identified GA segmentation is assumed to be large enough to be recognized as an actual GA region, the process may proceed to step M5.

[0057] Method step M5 analyzes the identified GA region by identifying one or two different metrics, each designed to identify one of two GA phenotypes, each associated with a high worsening rate GA. More specifically, the contour non-uniformity metric may be used to identify the "diffuse" phenotype GA, and the intensity uniformity metric may be used to identify the "striped" phenotype GA. When it is confirmed that the first identification metric is either the "diffuse" phenotype GA or the "striped" phenotype GA (e.g., the metric is higher than a predetermined threshold), method step M7 classifies the identified GA region as a "high worsening rate" GA. If it is not confirmed that the first identification metric is one of the "diffuse" or "striped" phenotypes, a second metric may be identified and checked to see if the other of the two phenotypes is present. If the other phenotype is present, the identified GA region may be classified again as a "high worsening rate" GA. Optionally, or alternatively, in method step M7, the specific phenotype classification ("diffuse" or "striped") identified for the GA region may be explicitly stated.

[0058] Fundus imaging system The two types of imaging systems used to capture fundus images are a flood illumination type imaging system (or flood illumination type imager) and a scanning illumination type imaging system (or scan imager). The flood illumination type imager simultaneously projects light onto the entire region of interest (FOV) of a specimen, for example using a flash lamp, and captures a full-frame image of the specimen (e.g., the fundus) with a full-frame camera (e.g., a camera having a two-dimensional (2D) photosensor array large enough to capture the entire desired FOV). For example, a flood illumination type fundus imager projects light onto the fundus and captures a full-frame image of the fundus in a single image capture sequence of the camera. The scan imager provides a scan beam that scans an object, e.g., an eye, and the scan beam is imaged at different scan positions while it scans the object, generating a series of image segments that may be reconstructed, e.g., montaged, to create a composite image of the desired FOV. The scan beam can be a point, a line, or a two-dimensional area such as a slit or a wide line.

[0059] FIG. 13 illustrates an example of a slit-scanning ophthalmic system SLO-1 for photographing an image of a fundus F, which is the inner surface of the eye E opposite the crystalline lens (i.e., the ocular lens) CL and may include the retina, optic nerve head, macula, fovea, and posterior pole. In this example, the imaging system is of a so-called "scan-descend" configuration, and the scanning line beam SB passes through the optical components of the eye E (including the cornea Crn, iris Irs, pupil Ppl, and ocular lens CL) to scan the fundus F. In the case of a flood-type fundus imager, a scanner is not required, and light is applied all at once over the entire desired field of view (FOV). Other scanning configurations are also known in the art, and the specific scanning configuration is not important for the present invention. As shown in the figure, the imaging system includes one or more light sources LtSrc, preferably a multi-color LED system or a laser system with appropriately adjusted etendue. An optional slit Slt (adjustable or stationary) may be positioned in front of the light source LtSrc and used to adjust the width of the scanning line beam SB. Additionally, the slit Slt may remain stationary during imaging or be adjusted to different widths to enable different confocal levels and any different uses during a particular scan or for reflection suppression. An optional objective lens ObjL may be installed in front of the slit Slt. The objective lens ObjL can be any one of the lenses according to the state of the art, including but not limited to refractive, diffractive, reflective, or hybrid lens / systems. The light from the slit Slt passes through the pupil splitting mirror SM and is directed to the scanner LnScn. It is desirable to bring the scan plane and the pupil plane as close as possible to reduce vignetting within the system. An optional optical system DL may be included to manipulate the optical distance between the images of the two components. The pupil splitting mirror SM may pass the illumination beam from the light source LtSrc to the scanner LnScn and reflect the detection beam (e.g., the reflected light returning from the eye E) from the scanner LnScn to the camera Cmr. The role of the pupil splitting mirror SM is to split the illumination and detection beams and assist in reflection suppression of the system.Scanner LnScn can also be a rotary galvo scanner or other types of scanners (e.g., piezoelectric, voice coil, microelectromechanical system (MEMS) scanners, electro-optic deflectors, and / or rotating polygon scanners). Depending on whether pupil splitting is performed before or after scanner LnScn, the scan can be divided into two steps, with one scanner in the illumination path and another scanner in the detection system. A particular pupil splitting device is described in detail in U.S. Patent No. 9,456,746, which is hereby incorporated by reference in its entirety.

[0060] From the scanner LnScn, the illumination beam passes through one or more optical systems, in this case the scanning lens SL and the ophthalmic or eye lens OL, thereby imaging the pupil of the eye F onto the image pupil of the system. Generally, the scanning lens SL receives scanning illumination beams at any of a plurality of scanning angles (angles of incidence) from the scanner LnScn and generates a scanning line beam SB on a focal plane of a substantially flat surface (e.g., a collimated optical path). The ophthalmic lens OL may focus the scanning line beam SB onto the fundus F (or retina) of the eye E to capture an image of the fundus. In this way, the scanning line beam SB generates a moving scanning line that moves through the fundus F. One possible configuration for these optical systems is a Keplerian telescope, in which case the distance between the two lenses is selected to create a substantially telecentric intermediate fundus image (4-f configuration). The ophthalmic lens OL can be a single lens, an achromatic lens, or a combination of different lenses. All lenses can be refractive, diffractive, reflective, or hybrid, as is known to those skilled in the art. The focal lengths of the ophthalmic lens OL and the scanning lens SL, as well as the size and / or shape of the pupil splitting mirror SM and the scanner LnScn, can be made different depending on the desired field of view (FOV). Thus, for example, by using a flip, an electric wheel, or a removable optical element within the optical system, a device can be envisioned that can switch a plurality of components to enter and exit the beam path according to the field of view. Since the size of the beam on the pupil varies with the change in the field of view, the pupil splitting can also be changed along with the change in the FOV. For example, a field of view of 45° to 60° is typical or standard for a fundus camera. A larger field of view, such as a wide-angle FOV of 60° to 120° or more, may also be achievable. The wide-angle FOV may be desirable for combinations with other imaging modalities such as a Broad-Line Fundus Imager (BLFI) or optical coherence tomography (OCT). The upper limit of the field of view may be determined by the physiological conditions around the human eye and the accessible working distance.Since the FOV of a typical human retina is 140° horizontally and 80° - 100° vertically, it may be desirable to have an asymmetric field of view for the highest FOV on the system.

[0061] The scanning line beam SB passes through the pupil Ppl of the eye E and is directed towards the retina, or fundus, surface F. The scanner LnScn1 adjusts the position of the light on the retina, or fundus F, such that a range of lateral positions on the eye E is illuminated. The reflected or scattered light (or emission in the case of fluorescence imaging) is directed back along a similar path as the illumination to define a collection beam CB on the detection path to the camera Cmr.

[0062] In the "scan - descan" configuration of this exemplary slit - scanning ophthalmic system SLO - 1, the light returning from the eye E is "descaned" by the scanner LnScn on the way to the pupil - splitting mirror SM. That is, the scanner LnScn scans the illumination beam from the pupil - splitting mirror SM to define the scanning illumination beam SB passing through the eye E. However, since the scanner LnScn also receives the light returning from the eye E at the same scan position, the scanner LnScn descans the returning light (e.g., cancels the scanning operation) and has the effect of defining a non - scanning (e.g., stable or stationary) condenser beam from the scanner LnScn to the pupil - splitting mirror SM, which bends the condenser beam towards the camera Cmr. At the pupil - splitting mirror SM, the reflected light (or emission in the case of fluorescence imaging) is separated from the illumination light to the detection path towards the camera Cmr, and the camera Cmr may be a digital camera having a photosensor for capturing an image. The imaging (e.g., objective) lens ImgL may be positioned within the detection path to image the fundus onto the camera Cmr. Similar to the case of the objective lens ObjL, the imaging lens ImgL may be of any type known in the art (e.g., refractive, diffractive, reflective, or hybrid lens). Details of other operations, especially methods for reducing image artifacts, are described in WO 2016 / 124644 pamphlet, the entire content of which is incorporated herein by reference. The camera Cmr captures the received image, for example, creates an image file, which can be further processed by one or more (electronic) processors or computing devices (e.g., the computer system shown in FIG. 20). Therefore, the condenser beam (returning from all scanning positions of the scanning line beam SB) may be condensed by the camera Cmr, and a full - frame image Img may be constructed from, for example, a montage of the individually captured condenser beams. However, other scanning configurations are also envisioned, including those where the illumination beam is scanned through the eye E and the condenser beam is scanned through the photosensor array of the camera.WO 2012 / 059236 pamphlet and US Patent Application Publication No. 2015 / 0131050 are incorporated herein by reference, and various designs are included where the return light is swept on the photo sensor array of the camera and where the return light is not swept on the photo sensor array of the camera.

[0063] In this example, camera Cmr is connected to a processor (e.g., processing module) Proc and a display (e.g., display module, computer screen, electronic screen, etc.) Dspl, both of which can be part of the image system itself or part of another dedicated processing and / or display unit, e.g., a computer system to which data is supplied from camera Cmr through a computer network including a cable or wireless network. The display and the processor can all be in one unit. The display can be a traditional electronic display / screen or a touch screen type, and can include a user interface for displaying information to the user and receiving information from the operator or user of the device. The user can interact with the display using any type of user input device known in the art, including but not limited to a mouse, knob, button, pointer, and touch screen.

[0064] During imaging, it may be desirable to keep the patient's line of sight fixed. One way to achieve this is to provide a fixation target that the patient's orientation can be changed to gaze at. The fixation target can be internal or external to the device depending on which part of the eye's image is to be generated. One embodiment of an internal fixation target is shown in FIG. 13. In addition to the first light source LtSrc used for imaging, one or more optional second light sources FxLtSrc, such as LEDs, can be positioned so that the light pattern is imaged on the retina using the lens FxL, the scanning element FxScn, and the reflector / mirror FxM. The fixation scanner FxScn can move the position of the light pattern, and the reflector FxM directs the light pattern from the fixation scanner FxScn towards the fundus F of the eye E. Preferably, the fixation scanner FxScn is positioned so that it is disposed in the pupil plane of the system, whereby the light pattern on the retina / fundus can be moved according to the desired fixation position.

[0065] The slit-scanning ophthalmoscope system can operate in different imaging modes depending on the light source and wavelength-selective filter element used. True-color reflection imaging (imaging similar to that observed when a clinician examines the eye using a hand-held or slit-lamp ophthalmoscope) can be achieved by taking an image of the eye with a series of color LEDs (red, blue, and green). The images of each color can be constructed step by step by turning on each LED at each scanning position, or each color image can be taken separately in its entirety. The three-color images can be combined to display a true-color image, or they can be displayed individually to highlight different features of the retina. The red channel highlights the choroid best, the green channel highlights the retina, and the blue channel highlights the pre-retinal layer. In addition, light of a specific frequency (e.g., individual color LEDs or lasers) can be used to excite different fluorophores in the eye (e.g., autofluorescence), and the resulting fluorescence can be detected by filtering out the excitation wavelength.

[0066] The fundus imaging system can also provide an infrared (IR) reflection image, for example, by using an infrared laser (or other infrared light source). The infrared (IR) mode is advantageous in that the eye is not sensitive to the IR wavelength. This allows the user to continuously capture images without disturbing the eye (e.g., during preview / alignment mode) to assist the user in aligning the device. Also, the IR wavelength may penetrate more into the tissue and better visualize the choroidal structure. In addition, imaging by fluorescence angiography (FA) and indocyanine green angiography (ICG) can be achieved by injecting a fluorescent dye into the bloodstream of the subject and then collecting the image.

[0067] Optical coherence tomography system In addition to fundus photographs, fundus autofluorescence (FAF), and fluorescence angiography (FA), ophthalmic images may be produced by other imaging modalities, such as optical coherence tomography (OCT), OCT angiography (OCTA), and / or ocular echography. The present invention, or at least a part thereof, may also be applied to these other ophthalmic imaging modalities with some modifications, as understood in the art. More specifically, the present invention may also be applied to ophthalmic images generated by an OCT / OCTA system that generates OCT and / or OCTA images. For example, the present invention may be applied to en face OCT / OCTA images. Examples of fundus images are provided in U.S. Patent Nos. 8,967,806 and 8,998,411, examples of OCT systems are provided in U.S. Patent Nos. 6,741,359 and 9,706,915, and examples of OCTA imaging systems may be found in U.S. Patent Nos. 9,700,206 and 9,759,544, the entire contents of all of which are incorporated herein by reference. For the sake of completeness, an example of an exemplary OCT / OCTA system is provided herein.

[0068] FIG. 14 illustrates a general-purpose frequency-domain optical coherence tomography (FD-OCT) system for collecting 3D image data of an eye suitable for use with the present invention. The FD-OCT system OCT_1 includes a light source LtSrc1. Typical light sources include, but are not limited to, broadband light sources with short temporal coherence lengths or swept laser sources. The beam of light from the light source LtScr1 is typically directed by an optical fiber Fbr1 to illuminate a sample, such as an eye E, and a typical sample is human intraocular tissue. The light source LrSrc1 can be either a broadband light source with a short temporal coherence length in the case of spectral-domain OCT (SD-OCT) or a wavelength-tunable laser source in the case of swept-source OCT (SS-OCT). The light is typically scanned by a scanner Scnr1 between the output of the optical fiber Fbr1 and the sample E, whereby the beam of light (dashed line Bm) scans the imaging target region of the sample in the lateral directions (x and y). In the case of full-field OCT, a scanner is not required and the light is applied to the entire desired field of view (FOV) at once. The light scattered from the sample is typically collected back into the same optical fiber Fbr1 that is used to guide the illumination light. The reference light derived from the same light source LtSrc1 travels along a separate path, which in this case includes an optical fiber Fbr2 and a retroreflector RR1 with an adjustable optical delay. As will be understood by those skilled in the art, a transmissive reference path can also be used and the adjustable delay can be placed within the sample or the reference arm of the interferometer. The collected sample light is typically combined with the reference light in a fiber coupler Cplr1 to form optical interference within an OCT detector Dtctr1 (e.g., a detector array, a digital camera, etc.). Although one fiber port is shown as reaching the detector Dtctr1, as will be understood by those skilled in the art, various designs of interferometers can be used for balanced or unbalanced detection of the interference signal. The output from the detector Dtctr1 is supplied to a processor Cmp1 (e.g., a computing device), which converts the observed interference into depth information of the sample. The depth information can be stored in a memory associated with the processor Cmp1 and / or displayed on a display (e.g., a computer / electronic display / screen) Scn1.The processing and storage functions may be located within the OCT device, or the functions may be performed on an external processing unit (e.g., the computer system shown in FIG. 20) to which the collected data is transferred. This unit may be dedicated to data processing or may be general and capable of performing other tasks that are not dedicated to the OCT device. The processor Cmp1 may include, for example, a field programmable gate array (FPGA), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a graphics processing unit (GPU), a system on chip (SoC), a central processing unit (CPU), a general purpose graphics processing unit (GPGPU), or a combination thereof that performs some or all of the data processing steps before or in parallel with being supplied to the host processor.

[0069] The sample and the reference arm within the interferometer can be configured with a bulk optical system, a fiber optic system, or a hybrid bulk optical system, and as is known to those skilled in the art, can have different architectures such as a Michelson, Mach-Zehnder, or common path system design. As used herein, a light beam should be construed as any optically directed path that is carefully directed. Instead of mechanically scanning the beam, the optical field can illuminate a one- or two-dimensional area of the retina to generate OCT data (see, for example, U.S. Patent No. 9,332,902, D. Hillmann et al., "Holoscopy-holographic optical coherence tomography," Optics Letters, Vol. 36(13), p. 2290, 2011, Y. Nakamura et al., "High-Speed three dimensional human retinal imaging by line field spectral domain optical coherence tomography," Optics Express, Vol. 15(12), p. 7103, 2007, Blazkiewicz et al., "Signal-to-noise ratio study of full-field Fourier-domain optical coherence tomography," Applied Optics, Vol. 44(36), p. 7722 (2005)). In a time domain system, the reference arm needs to have an adjustable optical delay to cause interference. A balanced detection system is typically used in TD-OCT and SS-OCT systems, and a spectrometer is used at the detection port for SD-OCT systems. The invention described herein can be applied to any type of OCT system.Various aspects of the present invention are applicable to any type of OCT system, or to a plurality of ophthalmic diagnostic systems including, but not limited to, other types of ophthalmic diagnostic systems and / or fundus imaging systems, perimetry devices, and scanning laser polarimeters.

[0070] In Fourier domain optical coherence tomography (FD-OCT), each measured value is a real-valued spectral control interference pattern (Sj(k)). Typically, several post-processing steps, including background removal, dispersion correction, etc., are performed on the real-valued spectral data. By Fourier transforming the processed interference pattern, a complex OCT signal output Aj(z)=|Aj|eiφ is obtained. From the absolute value of this complex OCT signal, |Aj|, the scattering intensity at different path lengths, and thus the profile of scattering with respect to depth (z-direction) within the sample, becomes apparent. Similarly, the phase φj can also be extracted from the complex OCT signal. The profile of backscattering with respect to depth is called an axial scan (A-scan). A cross-sectional image (tomographic image or B-scan) of the sample is generated by a set of A-scans measured at adjacent positions within the sample. A set of B-scans collected at different lateral positions on the sample constitutes a data volume or cube. For a particular data volume, the fast axis refers to the scan direction along one B-scan, and the slow axis refers to the axis along which multiple B-scans are collected. The term "cluster scan" may refer to one unit or block of data generated by repeated acquisitions at the same (or substantially the same) position (or region) in order to analyze the motion contrast that may be used to identify blood flow. A cluster scan can be composed of multiple A-scans or B-scans collected at relatively short time intervals at approximately the same position on the sample. Since the scans of a cluster scan are of the same region, the stationary structures remain relatively unchanged between scans during the cluster scan, while the motion contrast between scans meeting a certain criterion may be identified as blood flow. Various methods for generating B-scans are known in the art and include, but are not limited to, those along the horizontal or x-direction, those along the vertical or y-direction, those along the diagonal of x and y, or those of circular or spiral patterns. The B-scan may be within the x-z dimension, but may be any cross-sectional image including the z-dimension.

[0071] In OCT angiography or functional OCT, the analysis algorithm may be applied to OCT data collected at different times (e.g., cluster scans) at the same or substantially the same sample position on the sample in order to analyze motion or flow (e.g., see U.S. Patent Application Publication Nos. 2005 / 0171438, 2012 / 0307014, 2010 / 0027857, 2012 / 0277579, and U.S. Patent No. 6,549,801, the entireties of all of which are hereby incorporated by reference in their entirety). In an OCT system, any one of many OCT angiography processing algorithms (e.g., a motion contrast algorithm) may be used to identify blood flow. For example, a motion contrast algorithm can be applied to intensity information derived from the image data (intensity-based algorithm), phase information from the image data (phase-based algorithm), or complex image data (complex-based algorithm). An en face image is a 2D projection of 3D OCT data (e.g., by averaging the intensity of each individual A-scan, whereby each A-scan defines a pixel within the 2D projection). Similarly, an en face vasculature image is an image that displays a motion contrast signal, and the data dimension corresponding to depth therein (e.g., the z-direction along the A-scan) is typically displayed as one representative value (e.g., a pixel within the 2D projection image) by adding or integrating all or an isolated portion of the data (e.g., see U.S. Patent No. 7,301,644, the entirety of which is hereby incorporated by reference in its entirety). An OCT system that provides an angiography function may be referred to as an OCT angiography (OCTA) system.

[0072] FIG. 15 shows an example of an en face vasculature image. After processing the data and highlighting the motion contrast using any of the motion contrast methods known in the art, a pixel range corresponding to a certain tissue depth from the surface of the internal limiting membrane (ILM) of the retina may be added to generate an en face (e.g., front view) image of the vasculature.

[0073] Neural network As described above, the present invention may use a neural network (NN) machine learning (ML) model. To be thorough, a general overview of neural networks is provided herein. The invention may use any of the following neural network architectures, either alone or in combination. A neural network, or neural net, is a network of interconnected neurons (via nodes), where each neuron represents a node within the network. The set of neurons may be arranged in layers, and the output of one layer is fed forward to the next layer in a multi-layer perceptron (MLP) arrangement. An MLP may be understood as a feed-forward neural network that maps a set of input data to a set of output data.

[0074] FIG. 16 illustrates an example of a multi-layer perceptron (MLP) neural network. Its structure may include a plurality of hidden (e.g., inner) layers HL1-HLn, which map an input layer InL (which receives a set of inputs (or vector inputs) in_1-in_3) to an output layer OutL, which generates a set of outputs (or vector outputs), e.g., out_1 and out_2. Each layer may have any number of nodes, which are shown here as circles within each layer for illustrative purposes. In this example, the first hidden layer HL1 has two nodes, and the hidden layers HL2, HL3, and HLn each have three nodes. Generally, the deeper the MLP (e.g., the more hidden layers within the MLP), the greater its learning capacity. The input layer InL receives a vector input (shown here as a three-dimensional vector consisting of in_1, in_2, and in_3 for illustrative purposes) and may feed the received vector input to the first hidden layer HL1 within the sequence of hidden layers. The output layer OutL receives the output from the last hidden layer within the multi-layer model, e.g., HLn, and generates a vector output result (shown here as a two-dimensional vector consisting of out_1 and out_2 for illustrative purposes).

[0075] Typically, each neuron (i.e., node) generates one output, which is fed forward to the neurons in the subsequent layer. However, each neuron in the hidden layer may receive multiple inputs from the input layer or from the outputs of the neurons in the immediately preceding hidden layer. In general, each node may apply a function to its inputs to generate an output for that node. The nodes in the hidden layer (e.g., the learning layer) may apply the same function to their respective inputs to generate their respective outputs. However, some nodes, such as the nodes in the input layer InL, may receive only one input and may be passive, which means that they simply relay the value of that one input to their output. For example, they may provide a copy of their input to their output, which is indicated by the dashed arrows in the nodes of the input layer InL for illustrative purposes.

[0076] For the purpose of explanation, FIG. 17 shows a simplified neural network consisting of an input layer InL’, a hidden layer HL1’, and an output layer OutL’. The input layer InL’ is shown to have two input nodes i1 and i2, which receive Input_1 and Input_2 respectively (for example, the input nodes of layer InL’ receive a two-dimensional input vector). The input layer InL’ is fed forward to a single hidden layer HL1’ having two nodes h1 and h2, which in turn is fed forward to an output layer OutL’ of two nodes o1 and o2. The interconnections, or links, between neurons (shown as solid arrows for the purpose of explanation) have weights w1~w8. Typically, except for the input layer, a node (neuron) may receive as input the output of the nodes of the layer immediately preceding it. Each node multiplies each of its inputs by the corresponding interconnection weight of each input, adds the product of its inputs, and adds (or multiplies) a constant defined by other weights or biases that may be associated with that particular node (for example, node weights w9, w10, w11, and w12 corresponding to nodes h1, h2, o1, and o2 respectively), and then may calculate its output by applying a non-linear function or logarithmic function to the result. The non-linear function may be called an activation function or transfer function. Multiple activation functions are known in the art, and the choice of a particular activation function is not important for this explanation. However, it should be noted that the operations of the ML model and the behavior of the neural net depend on the values of the weights, which may be learned so that the neural network provides a desired output for a certain input.

[0077] During the training, or learning phase, a neural network learns (e.g., is trained to identify) appropriate weight values to achieve a desired output for a given input. Before the neural network is trained, each weight may be individually assigned an initial (e.g., random, optionally non-zero) value, such as a random number seed. Various methods for assigning initial weights are known in the art. Then, the weights are trained (optimized) such that for a given training vector input, the neural network produces an output close to the desired (predetermined) training vector output. For example, the weights may be gradually adjusted in thousands of repeated cycles by a method called backpropagation. In each cycle of backpropagation, the training input (e.g., vector input or training input image / sample) is passed through the neural network in a forward pass, providing its actual output (e.g., vector output). Then, the error for each output neuron, or output node, is calculated based on the actual neuron output and the teacher value training output for that neuron (e.g., the training output image / sample corresponding to the current training input image / sample). Then, it propagates backward through the neural network (from the output layer to the input layer in the reverse direction), and the weights are updated based on how much each weight contributes to the overall error, thereby bringing the output of the neural network closer to the desired training output. This cycle is then repeated until the actual output of the neural network is within an acceptable error range of the desired training output for its training input. As understood, each training input may require many backpropagation iterations to achieve the desired error range. Typically, an epoch refers to one backpropagation iteration (e.g., one forward pass and one backward pass) over all training samples, and many epochs may be required to train a neural network. Generally, the larger the training set, the better the performance of the trained ML model, so various data augmentation methods may be used to increase the size of the training set.For example, if the training set includes pairs of training input images and training output images to which it corresponds, the training images may be divided into a plurality of corresponding image segments (or patches). Corresponding patches from the training input image and the training output image are paired, and a plurality of training patch pairs may be defined from one input / output image pair, thereby expanding the training set. However, training a large training set increases the requirements for computer resources, such as memory and data processing resources. The computational requirements may be reduced by dividing the large training set into a plurality of mini-batches, and the size of this mini-batch is determined by the number of training samples in one forward / backward pass. In this case, and one epoch may include a plurality of mini-batches. Another problem is that the NN may overfit the training set, reducing its ability to generalize from specific inputs to different inputs. The problem of overfitting may be reduced by creating an ensemble of neural networks or by randomly dropping out nodes within the neural network during training, which effectively removes the dropped leads from the neural network. Various dropout adjustment methods, such as inverse dropout, are known in the art.

[0078] It should be noted that the operations of a trained NN machine model are not simple algorithms for the operation / analysis step. In fact, when a trained NN machine model receives an input, the input is not analyzed in the conventional sense. Rather, regardless of the gist and nature of the input (e.g., a vector defining a live image / scan, or a vector defining any other entity such as a demographic description or a record of activities), the input is subject to the same architectural construction of the trained neural network (e.g., the same node / layer arrangement, trained weight and bias values, a given convolution / transposed convolution operation, activation function, pooling operation, etc.), and it may not be clear how the architectural construction of the trained network generates its output. Furthermore, the values of the trained weights and biases are not deterministic and depend on many factors such as the amount of time for training given to the neural network (e.g., the number of epochs in training), the random starting values of the weights before training begins, the computer architecture of the machine on which the NN is trained, the selection of training samples, the distribution of training samples among multiple mini-batches, the selection of activation functions, the selection of the error function for changing the weights, and even whether the training is interrupted on one machine (e.g., having a first computer architecture) and completed on another machine (e.g., having a different computer architecture). The point is that it is not clear why a trained ML model reaches a particular output, and many studies are currently being conducted to identify the elements on which the ML model bases its output. Therefore, the processing of neural networks for live data cannot be reduced to a simple step algorithm. Rather, the operations depend on its training architecture, training sample set, training sequence, and various circumstances in the training of the ML model.

[0079] Broadly speaking, the configuration of an NN machine learning model may include a learning (or training) stage and a classification (or inference) stage. In the learning stage, the neural network may be trained for a specific purpose, and a set of training examples may be provided, which includes training (sample) inputs and training (sample) outputs, and optionally, a set of validation examples for testing the progress of training. During this learning process, various weights associated with the nodes and node interconnections within the neural network are gradually adjusted to reduce the error between the actual output of the neural network and the desired training output. In this way, a multi-layer feedforward neural network (such as the one described above) may be able to approximate any measurable function to any desired accuracy. The result of the learning stage is a learned (e.g., trained) (neural network) machine learning (ML). In the inference stage, a set of test inputs (or live inputs) may be provided to the learned (trained) ML model, which may apply what it has learned to generate an output prediction based on the test inputs.

[0080] Similar to the ordinary neural networks of FIGS. 15 and 16, a convolutional neural network (CNN) is also composed of neurons with learnable weights and biases. Each neuron receives an input, performs an operation (e.g., dot product), and optionally followed by a non-linear transformation. However, a CNN receives raw image pixels at one end (e.g., the input end) and provides classification (or class) scores at the opposite end (e.g., the output end). Since a CNN expects images as inputs, these are optimized to handle volumes (e.g., the pixel height and width of the image and the color depth, such as the RGB depth defined by three colors: red, green, and blue). For example, the layers of a CNN may be optimized for neurons arranged in three dimensions. Neurons within a CNN layer may be connected to a small region of the previous layer, rather than to all of the neurons of a fully connected NN. The final output layer of a CNN may reduce the full image to one vector (classification) arranged along the depth dimension.

[0081] FIG. 18 provides an exemplary convolutional neural network architecture. The convolutional neural network may be defined as a succession of two or more layers (e.g., layers 1 through N), where the layers may include an (image) convolution step, a weighted sum step (of the result), and a non-linear function step. Convolution may be performed on its input data, for example, by applying a filter (or kernel) over a moving window across the input data to generate a feature map. Each layer and layer component may have different predetermined filters (from a filter bank), weights (or weighting parameters), and / or function parameters. In this example, the input data is an image of a certain pixel height and width, which may be the raw pixel values of the image. In this example, the input image is depicted as having a depth of three color channels RGB (red, green, blue). Optionally, various preprocessing may be performed on the input image, and the result of the preprocessing may be input instead of, or in addition to, the raw image data. Some examples of image processing may include retinal blood vessel map segmentation, color space conversion, adaptive histogram equalization, connected component generation, etc. Within a layer, a dot product may be calculated between certain weights and small regions to which they are connected within the input volume. Many ways of constructing a CNN are known in the art, but by way of example, the layers may be configured to apply an element-wise activation function such as the max(0,x) threshold at zero. A pooling function may be performed (e.g., along the x-y directions) to downsample the volume. A fully connected layer may be used to identify a classification output and generate a one-dimensional output vector that has been found to be useful for image recognition and classification. However, for image segmentation, the CNN needs to classify each pixel. Since each CNN layer has a tendency to reduce the resolution of the input image, another stage is needed to upsample the image back to its original resolution. This may be achieved by applying a transposed convolution (or deconvolution) stage TC, which typically has learnable parameters instead of using any predetermined interpolation method.

[0082] Convolutional neural networks have been successfully applied to many problems in computer vision. As mentioned above, generally, a large training dataset is required to train a CNN. The U-Net architecture is based on CNN and can generally be trained with a smaller training dataset than conventional CNNs.

[0083] FIG. 19 illustrates an exemplary U-Net architecture. This exemplary U-Net includes an input module (or input layer or stage), which receives an input U-in (e.g., an input image or image patch) of any size (e.g., a size of 128×128 pixels). The input image may be a fundus image, an OCT / OCTA en face, a B-scan image, etc. However, it should be understood that the input may be of any size and dimension. For example, the input image may be an RGB color image, a monochrome image, a volume image, etc. The input image passes through a series of processing layers, each of which is illustrated with an exemplary size, but these sizes are for illustrative purposes only and will depend, for example, on the size of the image, the convolutional filter, and / or the pooling stage. This architecture consists of a contracting path (including four encoding modules), followed by an expanding path (including four decoding modules), and four copy-and-crop links (e.g., CC1 to CC4) that copy the output of one encoding module within the contracting path and combine it with the input of the corresponding decoding module within the expanding path. As a result, a characteristic U-shape is formed, from which this architecture gets its name. The contracting path is similar to an encoder, and its basic function is to capture context through a compact feature map. In this example, each encoding module within the contracting path includes two convolutional neural network layers, followed by an optional max pooling layer (e.g., a downsampling layer). For example, the input image U_in passes through two convolutional layers, each having 32 feature maps. The number of feature maps doubles at each pooling, starting from 32 feature maps in the first block and becoming 64 in the second block, etc. Therefore, the contracting path forms a convolutional network consisting of a plurality of encoding modules (or stages), each providing a convolutional stage, followed by an activation function (e.g., a rectified linear unit ReLU or a sigmoid layer) and a max pooling operation.The expansion path is similar to the decoder, and its function is to provide localization and retain spatial information regardless of the downsampling performed in the contraction stage and any max-pooling. In the convergence path, the spatial information is reduced and the feature information is increased. The expansion path includes a plurality of decoding modules, and each decoding module combines its current value with the output of the corresponding encoding module. That is, the features and spatial information are combined through a sequence of up-convolutions (e.g., upsampling or transposed convolution, i.e., deconvolution) in the expansion path and the combination (e.g., via CC1~CC4) with the high-resolution features from the convergence path. Therefore, the output of the deconvolution layer is combined with the corresponding (optionally cropped) feature map from the convergence path, followed by two convolutional layers and an activation function (batch normalization optionally). The output from the last module within the expansion path may be fed to other processing / training blocks or layers such as a classifier block, which may be trained together with the U-Net architecture.

[0084] The module / stage (BN) between the convergence path and the expansion path may be referred to as a "bottleneck". The bottleneck BN may consist of two convolutional layers (with batch normalization and dropout optionally).

[0085] Computing device / system FIG. 20 illustrates an exemplary computer system (or computing device or computer device). In some embodiments, one or more computer systems may provide the functions described or illustrated herein and / or perform one or more steps of one or more of the methods described or illustrated herein. The computer system may take any suitable physical form. For example, the computer system may be an embedded computer system, a system-on-chip (SOC), or a single-board computer system (SBC) (e.g., a computer-on-module (COM) or a system-on-module (SOM), etc.), a desktop computer system, a laptop or notebook computer system, a mesh of computer systems, a mobile phone, a personal digital assistant (PDA), a server, a tablet computer system, an augmented / virtual reality device, or a combination of two or more of these. Optionally, the computer system may be in the cloud, which may include one or more cloud components within one or more networks.

[0086] In some embodiments, the computer system may include a processor Cpnt1, a memory Cpnt2, a storage Cpnt3, an input / output (I / O) interface Cpnt4, a communication interface Cpnt5, and a bus Cpnt6. Optionally, the computer system may also include a display Cpnt7, such as a computer monitor or screen.

[0087] Processor Cpnt1 includes hardware for executing instructions, such as those that make up a computer program. For example, Processor Cpnt1 may be a central processing unit (CPU) or a general-purpose computing on graphics processing unit (GPGPU). Processor Cpnt1 may read (or fetch) instructions from an internal register, internal cache, Memory Cpnt2, or Storage Cpnt3, decode and execute this instruction, and write one or more results to an internal register, internal cache, Memory Cpnt2, or Storage Cpnt3. In certain embodiments, Processor Cpnt1 may include one or more internal caches for data, instructions, or addresses. Processor Cpnt1 may include one or more instruction caches, one or more data caches, for example, to hold data tables. Instructions in the instruction cache may be copies of instructions in Memory Cpnt2 or Storage Cpnt3, and the instruction cache may speed up the reading of these instructions by Processor Cpnt1. Processor Cpnt1 may include any suitable number of internal registers and may include one or more arithmetic logic units (ALUs). Processor Cpnt1 may be a multi-core processor or may include one or more Processor Cpnt1s. Although the present disclosure describes and illustrates a particular processor, the present disclosure contemplates any suitable processor.

[0088] Memory Cpnt2 may include a main memory that stores instructions for the processor Cpnt1 that executes processes or holds intermediate data during processing. For example, a computer system may load instructions or data (e.g., a data table) from storage Cpnt3 or from another source (e.g., another computer system) into memory Cpnt2. Processor Cpnt1 may load instructions and data from memory Cpnt2 into one or more internal registers or an internal cache. To execute an instruction, processor Cpnt1 may read and decode the instruction from the internal register or internal cache. During or after execution of an instruction, processor Cpnt1 may write one or more results (which may be intermediate or final results) to an internal register, internal cache, memory Cpnt2, or storage Cpnt3. Bus Cpnt6 may include one or more memory buses (each of which may include an address bus and a data bus), and may couple processor Cpnt1 to memory Cpnt2 and / or storage Cpnt3. Optionally, one or more memory management units (MMUs) facilitate data transmission between processor Cpnt1 and memory Cpnt2. Memory Cpnt2 (which may be a high-speed volatile memory) may include random access memory (RAM), such as dynamic RAM (DRAM) or static RAM (SRAM). Storage Cpnt3 may include long-term or high-capacity storage for data or instructions. Storage Cpnt3 may be built into or external to the computer system, and may include one or more of a disk drive (e.g., a hard disk drive HDD or a solid state drive SSD), flash memory, ROM, EPROM, optical disk, magneto-optical disk, magnetic tape, universal serial bus (USB)-accessible drive, or other types of non-volatile memory.

[0089] The I / O interface Cpnt4 may be software, hardware, or a combination of both, and may include one or more interfaces (e.g., serial or parallel communication ports) for communicating with an I / O device, which may enable communication with a human (e.g., a user). For example, I / O devices may include a keyboard, keypad, microphone, monitor, mouse, printer, scanner, speaker, steel camera, stylus, table, touch screen, trackball, video camera, other suitable I / O devices, or a combination of two or more of these.

[0090] The communication interface Cpnt5 may provide a network interface for communicating with other systems or networks. The communication interface Cpnt5 may include a Bluetooth® interface or other types of packet-based communication. For example, the communication interface Cpnt5 may include a network interface controller (NIC) and / or a wireless NIC or wireless adapter for communication with a wireless network. The communication interface Cpnt5 may provide communication with a WI-FI network, ad hoc network, personal area network (PAN), wireless PAN (e.g., Bluetooth WPAN), local area network (LAN), wide area network (WAN), metropolitan area network (MAN), cellular phone network (e.g., Global System for Mobile Communications (GSM®) network, etc.), the Internet, or a combination of two or more of these.

[0091] Bus Cpnt6 may provide a communication link between the above-described components of the computing system. For example, Bus Cpnt6 may be an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand bus, a low-pin-count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCIe) bus, a serial advanced technology attachment (SATA) bus, a Video Electronics Standards Association local (VLB) bus, or other suitable bus, or a combination of two or more of these.

[0092] The present disclosure describes and illustrates a particular computer system having a particular number of particular components in a particular arrangement, but the present disclosure contemplates any suitable computer system having any suitable number of any suitable components in any suitable arrangement.

[0093] As used herein, a computer-readable non-transitory storage medium may include one or more semiconductor-based or other integrated circuits (ICs) (e.g., field programmable gate arrays (FPGAs) or application specific ICs (ASICs)), hard disk drives (HDDs), hybrid hard drives (HHDs), optical disks, optical disk drives (ODDs), magneto-optical disks, magneto-optical drives, floppy diskettes, floppy disk drives (FDDs), magnetic tapes, solid state drives (SSDs), RAM-drives, SECURE DIGITAL cards or drives, any other suitable computer-readable non-transitory storage medium, or any suitable combination of two or more thereof, as appropriate. A computer-readable non-transitory storage medium may be volatile, non-volatile, or a combination of volatile and non-volatile, as appropriate.

[0094] While the present invention has been described in conjunction with several specific embodiments, as will be apparent to those skilled in the art in light of the foregoing description, many other alternatives, modifications, and variations will be apparent. Accordingly, the invention described herein is intended to embrace all such alternatives, modifications, applications, and variations that may fall within the spirit and scope of the appended claims.

Claims

1. A method of operating a computer processor to classify geographic atrophy (GA) in the eye, comprising: obtaining an image of the fundus of an eye using an ophthalmic diagnostic device; wherein the computer processor providing the obtained image as an input to a phenotypic classifier based on a machine learning model, the phenotypic classifier being a neural network trained using a set of training output images each showing the phenotype of a plurality of GA regions identified by an expert as a training output set, and the trained phenotypic classifier identifying GA regions from the input image and generating, in an output, a result of classifying the identified GA regions into a striate phenotype or a diffuse phenotype; displaying the result of the phenotypic classifier on an electronic display or storing the result for further processing.

2. The phenotypic classifier is a neural network trained using a training output set of fundus images having manually delineated GA regions whose phenotypes have been identified, and a training input set of fundus images of the same eye that do not have delineated GA regions and do not have identified phenotypes; The method according to claim 1, wherein retinal blood vessels are removed from the fundus images of the training input set and the training output set before training the neural network.

3. further comprising identifying a region of interest (ROI) within the obtained image, the ROI including retinal maculopathy of the eye; The method according to claim 1 or 2, wherein the phenotypic classifier limits processing to the ROI.

4. further comprising identifying an image quality (IQ) measure of the obtained image; The method according to any one of claims 1 to 3, wherein the obtained image is provided to the phenotypic classifier in response to the measured identified image quality being higher than a predetermined IQ threshold.

5. The phenotypic classification is based on a junction zone of the identified GA regions.

6. For each identified GA region, the phenotypic classifier maps the currently identified GA region into 3D space and identifies the surface area of the 3D mapped GA region. The method according to any one of claims 1 to 5, wherein in response to the surface area of the 3D mapped GA region being smaller than a predetermined minimum area threshold, the currently identified GA region is reclassified as a non-GA region.

7. For each identified GA region, the phenotype classifier identifies a first measure, which is one of a contour non-uniformity measure or an intensity non-uniformity measure that is adjacent to and along the perimeter of the currently identified GA region, and in response to the identified first measure being higher than a first threshold, classifies the currently identified GA region as a high progression rate GA. The method according to any one of claims 1 to 6.

8. The phenotype classifier in response to the first measure being less than or equal to the first threshold, identifies a second measure, which is the other of the contour non-uniformity measure or the intensity non-uniformity measure, and in response to the second measure being higher than a second threshold, classifies the currently identified GA region as a high progression rate GA. The method according to claim 7.

9. For the currently identified GA region, the phenotype classifier identifies a contour non-uniformity measure that is adjacent to and along the perimeter of the currently identified GA region, and in response to the identified contour non-uniformity measure being higher than a predetermined non-uniformity threshold, classifies the currently identified GA region as a "diffuse" phenotype. The method according to any one of claims 1 to 8.

10. The method according to claim 9, wherein the non-uniformity threshold is based on the standard mean contour deviation of a library of sample GA regions.

11. For the currently identified GA region, the phenotype classifier identifies an intensity non-uniformity measure that is adjacent to and along the perimeter of the currently identified GA region, and in response to the identified intensity non-uniformity measure being higher than a predetermined intensity non-uniformity threshold, classifies the currently identified GA region as a "striped" phenotype. The method according to any one of claims 1 to 10.

12. Identifying the intensity non-uniformity measure includes selecting a set of random points along the perimeter of the currently identified GA region and identifying peaks and valleys of pixel intensity along a direction perpendicular to the perimeter extending from the random points. The method according to claim 11, wherein the intensity non-uniformity threshold is based on the normalized mean of peaks and valleys of pixel intensity along the junction zone of sample GA regions.

13. The method according to any one of claims 1 to 12, wherein retinal blood vessels are removed from the acquired image before identifying the GA region.

14. A method of operating a computer processor to classify geographic atrophy (GA) in an eye, comprising: acquiring an image of the fundus of the eye using an ophthalmic diagnostic device; providing, by the computer processor, the image to an automatic GA identification process that identifies GA regions in the image; the computer processor identifying a first metric, the first metric being one of a contour non-uniformity metric or an intensity non-uniformity metric in a region of the image adjacent to and along the perimeter of the identified GA region; the computer processor generating a result of classifying the GA region as high-progression GA if the identified first metric is higher than a first threshold; the computer processor displaying the classified result of the GA region on an electronic display or storing the classified result for further processing.

15. The computer processor identifies a second metric if the first metric is less than or equal to the first threshold, the second metric being the remaining one of the contour non-uniformity metric or the intensity non-uniformity metric; identifying the currently identified GA region as high-progression GA if the second metric is higher than a second threshold, the method according to claim 14.

16. The computer processor identifies a region of interest (ROI) in the image of the fundus based at least on a predetermined anatomical landmark using an automatic ROI identification process, the ROI including the retinal macula of the eye; the automatic GA identification process being limited to image data within the ROI, the method according to claim 14 or 15.

17. The automatic region of interest (ROI) identification process identifying the location of the major vascular density in the image; identifying the optic nerve head at least partially based on the identified location; using the identified optic nerve head and identification of whether the image is of the left or right eye as a reference for fovea search; defining the ROI based on the location of the fovea, the method according to claim 16.

18. The ROI includes a plurality of sub-image sectors, and the automatic GA identification process classifies each sub-image sector of the ROI as a GA sub-image or a non-GA sub-image for each sub-image sector. The method according to claim 16 or 17.

19. The method according to claim 18, wherein each sub-image sector includes one pixel.

20. The method according to any one of claims 14 to 19, wherein retinal blood vessels are removed from the image before being provided to the automatic GA identification process.

21. The method according to any one of claims 14 to 20, wherein the automatic GA identification process includes providing the image to a trained machine learning (ML) model that provides an initial GA segmentation.

22. The method according to claim 21, wherein the ML model is a support vector machine.

23. The method according to claim 21, wherein the ML model is a trained neural network based on a U-Net architecture.

24. The automatic GA identification process includes providing the result of the ML model to a dynamic contour algorithm that uses the initial GA segmentation as a starting point, and the result of the dynamic contour algorithm at least partially identifies GA lesions in the image. The method according to any one of claims 21 to 23.

25. The method according to claim 24, wherein the dynamic contour algorithm is based on Chan-Vese segmentation improved to change the energy and movement direction of contour growth.

26. The image is a planar image including two-dimensional (2D) pixels, and the automatic GA identification process includes mapping the identified GA region into a three-dimensional (3D) space to define a 3D mapped GA region, identifying the area of the 3D mapped GA region, and re-identifying the GA region as a non-GA region when the identified area of the 3D mapped GA region is smaller than a predetermined area threshold. The method according to any one of claims 14 to 25.

27. Identifying the contour non-uniformity measure includes selecting a set of random points along the perimeter of the identified GA region, and identifying the distance from each random point to the centroid of the identified GA region. The method according to any one of claims 14 to 26, wherein the contour non-uniformity measure is based on the variability of the identified distances.

28. The method according to claim 27, wherein the set of random points is evenly distributed at equal distances along the periphery of the GA region.

29. The method according to claim 27 or 28, wherein the identified GA region is classified into a "diffuse" phenotype based on the contour non-uniformity measure.

30. Identifying the intensity uniformity measure involves selecting a set of random points along the periphery of the identified GA region, and identifying the intensity variability measure for each random point based on the peaks and valleys of pixel intensities along a direction perpendicular to the surrounding contour extending from each random point, and the intensity uniformity measure is based on the intensity variability measure. The method according to any one of claims 14 to 29, wherein the intensity uniformity measure is based on the intensity variability measure.

31. The method according to claim 30, wherein the identified GA region is classified into a "striped" phenotype when at least a predetermined percentage of the intensity variability measure is greater than a predetermined variability threshold.

32. The method according to claim 30 or 31, wherein identifying the intensity variability measure includes using at least one of a Hessian filter gradient derivative or a directional Gaussian filter.

33. A system for classifying geographic atrophy (GA) in the eye, comprising an ophthalmic diagnostic device for acquiring an image of the fundus of the eye, a computer processor, a phenotype classifier based on a machine learning model, and a display and / or storage means, and is configured to execute the method according to any one of claims 1 to 13.

34. A system for classifying geographic atrophy (GA) in the eye, comprising an ophthalmic diagnostic device for acquiring an image of the fundus of the eye, a computer processor, an automatic GA identification processor for identifying GA regions in the image, and a display and / or storage means, and is configured to execute the method according to any one of claims 14 to 32. ​ ​

Citation Information

Patent Citations

  • Geographic atrophy identification and measurement

    JP2016131881A

  • Optic disc detection in retinal autofluorescence images

    US20160345819A1

  • Generalizable medical image analysis using segmentation and classification neural networks

    US20190005684A1