Multimodal geographic atrophic lesion segmentation

A multimodal imaging approach using FAF, NIR, and OCT images with neural networks addresses the challenges of GA segmentation, achieving precise and efficient lesion quantification and monitoring.

JP7853292B2Active Publication Date: 2026-04-28GENENTECH INC
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
GENENTECH INC
Filing Date
2021-10-22
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Current methods for evaluating geographic atrophy (GA) in the retina, characterized by the progressive loss of photoreceptors and retinal pigment epithelium, face challenges such as manual segmentation being time-consuming and prone to variability, and automated methods struggling with low intensity regions like the fovea, leading to inaccurate lesion quantification.

Method used

A multimodal approach using fundus autofluorescence (FAF) images combined with near-infrared (NIR) and optical coherence tomography (OCT) images, employing neural networks like U-Net and Y-Net for automated GA lesion segmentation, enhancing accuracy and efficiency.

Benefits of technology

The method provides accurate and efficient segmentation of GA lesions, enabling improved long-term monitoring and quantification of GA progression, reducing inter- and intra-observer variability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007853292000003
    Figure 0007853292000003
  • Figure 0007853292000004
    Figure 0007853292000004
  • Figure 0007853292000005
    Figure 0007853292000005
Patent Text Reader

Abstract

Disclosed herein are methods and systems for generating a geographic atrophy (GA) lesion segmentation mask corresponding to one or more retinal GA lesions. In some embodiments, a set of fundus autofluorescence (FAF) images of a retina having one or more GA lesions and one or both of a set of infrared (IR) images of the retina or a set of optical coherence tomography (OCT) images of the retina can be used to generate a GA lesion segmentation mask including one or more GA lesion segments corresponding to one or more retinal GA lesions. In some cases, a neural network can be used to generate the GA lesion segmentation mask.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Cross-reference of related applications This application claims priority and interest to U.S. Provisional Patent Application No. 63 / 105,105, filed on 23 October 2020, and U.S. Provisional Patent Application No. 63 / 218,908, filed on 6 July 2021, both entitled “Multimodal Geographic Atrophy Lesion Segmentation,” which are incorporated herein by reference in their entirety as described below and for all applicable purposes.

[0002] field This description relates in general to the evaluation of geographic atrophy in the retina. More specifically, this description provides a method and system for evaluating geographic atrophy using images from multiple modalities, including fundus autofluorescence (FAF) images, as well as one or both of near-infrared (NIR) and optical coherence tomography (OCT) images. [Background technology]

[0003] background Age-related macular degeneration (AMD) is the leading cause of vision loss in patients over 50 years of age. Geographic atrophy (GA) is one of the two progressive stages of AMD and is characterized by the progressive and irreversible loss of choroidal capillaries, retinal pigment epithelium (RPE), and photoreceptors. Diagnosis and monitoring of GA lesion expansion can be performed using fundus autofluorescence (FAF) images obtained by confocal scanning laser ophthalmography (cSLO). This type of imaging technique, which shows the topographic mapping of lipofuscin in the RPE, can be used to measure changes in GA lesions over time. In FAF images, GA lesions appear as clearly defined areas of low autofluorescence, resulting from the loss of RPE, and therefore lipofuscin. However, quantifying GA lesions on FAF images can be difficult due to the naturally occurring low intensity in the fovea. Furthermore, quantifying GA lesions using FAF images is typically a manual process that is more time-consuming than desired and prone to inter- and intra-observer variability. Therefore, it may be desirable to have one or more methods, systems, or both that recognize and take into consideration one or more of these problems. [Overview of the Initiative]

[0004] overview Some embodiments of the present disclosure include a method comprising receiving a set of fundus autofluorescence (FAF) images of a retina having one or more geographic atrophy (GA) lesions. The method further includes receiving one or both of a set of infrared (IR) images of the retina or a set of optical coherence tomography (OCT) images of the retina, and using the set of FAF images and one or both of the set of IR images or OCT images to generate a GA lesion segmentation mask containing one or more GA lesion segments corresponding to one or more GA lesions in the retina.

[0005] In some embodiments, the system comprises non-temporary memory and a hardware processor coupled to the non-temporary memory and configured to read instructions from the non-temporary memory and cause the system to perform an operation. In some cases, the operation includes receiving a set of FAF images of the retina having one or more GA lesions, receiving one or both of a set of IR images of the retina or a set of OCT images of the retina, and using the set of FAF images and one or both of the set of IR images or OCT images to generate a GA lesion segmentation mask containing one or more GA lesion segments corresponding to one or more GA lesions in the retina.

[0006] Some embodiments of the present disclosure disclose a non-temporary computer-readable medium (CRM) storing executable computer-readable instructions causing a computer system to perform an operation including receiving a set of FAF images of a retina having one or more GA lesions. In some embodiments, the operation further includes receiving one or both of a set of IR images of the retina or a set of OCT images of the retina, and using the set of FAF images and one or both of the set of IR images or OCT images to generate a GA lesion segmentation mask containing one or more GA lesion segments corresponding to one or more GA lesions in the retina.

[0007] Other aspects, features, and embodiments of the present invention will become apparent to those skilled in the art by examining the following description of a particular exemplary embodiment of the present invention in conjunction with the accompanying drawings. While the features of the present invention can be described with respect to the following specific embodiments and drawings, all embodiments of the present invention may include one or more of the advantageous features described herein. In other words, one or more embodiments may be described as having a particular advantageous feature, but one or more such features may also be used according to the various embodiments of the present invention described herein. Similarly, while exemplary embodiments may be described below as embodiments of an apparatus, system, or method, it should be understood that such exemplary embodiments may be implemented in a variety of systems and methods. [Brief explanation of the drawing]

[0008] For a more complete understanding of the principles and advantages disclosed herein, refer to the following description in conjunction with the accompanying drawings.

[0009] [Figure 1] This is a block diagram of a lesion evaluation system according to various embodiments.

[0010] [Figure 2] This is a flowchart of the process for evaluating geographic atrophy in various embodiments.

[0011] [Figure 3A-3B] This paper demonstrates the use of the U-Net deep learning neural network for evaluating geographic atrophic lesions in various embodiments.

[0012] [Figure 4A-4B] This paper demonstrates the Y-Net deep learning neural network and its use for evaluating geographic atrophic lesions in various embodiments.

[0013] [Figure 5] An exemplary workflow for segmenting FAF images and NIR images using deep learning neural networks of Y-Net and U-Net according to various embodiments is shown.

[0014] [Figure 6A-6B] Exemplary segmentation results of FAF images and NIR images using deep learning neural networks of Y-Net and U-Net according to various embodiments are shown.

[0015] [Figure 7] An exemplary Dice similarity coefficient score for measuring the similarity between segmentations performed by a Y-Net deep learning neural network, a U-Net deep learning neural network, and a human evaluator according to various embodiments is shown.

[0016] [Figures 8A-8D] An exemplary comparison of lesion area sizes at screening measured by a Y-Net deep learning neural network, a U-Net deep learning neural network, and a human evaluator according to various embodiments is shown.

[0017] [Figures 9A-9E] An exemplary comparison of lesion area sizes at the 12th month measured by a Y-Net deep learning neural network, a U-Net deep learning neural network, and a human evaluator according to various embodiments is shown.

[0018] [Figure 10] An exemplary neural network that can be used to implement a Y-Net deep learning neural network and a U-Net deep learning neural network according to various embodiments is shown.

[0019] [Figure 11]This is a block diagram of a computer system according to various embodiments.

[0020] It should be understood that the drawings are not necessarily drawn to a consistent scale, and the objects within the drawings are not necessarily drawn to a consistent scale with respect to each other. The drawings are intended to provide clarity and understanding of the various embodiments of the apparatus, systems, and methods disclosed herein. Wherever possible, the same reference numerals are used throughout the drawings to refer to the same or similar parts. Furthermore, it should be understood that the drawings are not in any way limited to the scope of this instruction. [Modes for carrying out the invention]

[0021] Detailed explanation I. Overview Current methods for evaluating geographic atrophy (GA), characterized by the progressive and irreversible loss of photoreceptors, retinal pigment epithelium (RPE), and choroidal capillary tubules, involve analyzing fundus autofluorescence (FAF) images to assess GA lesions, which are detected and bordered due to the reduction in autofluorescence caused by the loss of RPE cells and the endogenous fluorophore lipofuscin. Depiction or segmentation of GA lesions involves creating a pixel-level mask of GA lesions within the FAF image. The pixel-level mask identifies each pixel as belonging to at least one of two distinct classes. As an example, each pixel can be assigned to either a first class corresponding to GA lesions or a second class not corresponding to GA lesions. In this way, pixels assigned to the first class identify GA lesions. This type of segmentation is sometimes called GA segmentation.

[0022] Segmentation of GA lesions in FAF images can be manual or automated. Manual measurement of GA lesions from FAF images by human evaluators can provide excellent reproducibility if performed by experienced evaluators, but manual depiction of GA regions is time-consuming and susceptible to inter- and intra-evaluator variability, as well as variability between different reading centers, especially with less experienced evaluators. Several methods involve using software that can be semi-automated to segment GA lesions in FAF images. However, for high-quality segmentation, a trained reader of FAF images may still be necessary, for example, using human user input to accurately depict lesion boundaries. Therefore, there is a need for methods and systems that enable fully automated segmentation of GA lesions in FAF images in the retina.

[0023] Artificial intelligence (AI), particularly machine learning and deep learning systems, can be configured to automatically segment GA lesions or to evaluate the shape-descriptive features of GA progression from retinal FAF images. AI systems can detect and quantify GA lesions at least as accurately as human raters, especially when processing large datasets, but much faster and more cost-effectively. For example, segmentation performed by algorithms using k-nearest neighbors (k-NN) pixel classifiers, fuzzy c-means (clustering algorithms), or deep convolutional neural networks (CNNs) can have good agreement with manual segmentation performed by trained raters. That said, FAF images tend to show low intensity in the portion of the image that shows the fovea of ​​the subject's retina. Low intensity means that the boundaries or regions of lesions may be indistinct or difficult to distinguish. As a result, accurately quantifying GA lesions from FAF images is difficult, even with AI or manual annotation.

[0024] Simultaneously, various other imaging modalities may be used in clinical trials or clinical practice to capture retinal images. For example, infrared reflectivity (IR) imaging and optical coherence tomography (OCT) imaging are other common imaging methods. Currently, imaging modalities are used independently and mutually exclusive of each other when evaluating the progression of GA.

[0025] This embodiment aims to leverage the beneficial properties of retinal images provided by various imaging modalities in efforts to improve the evaluation of GA progression. GA lesion segmentation results obtained from retinal images, which are a combination of various imaging techniques, are more accurate and can provide additional information compared to segmentation results obtained from unimodal image inputs. Therefore, the methods and systems of this disclosure enable automated segmentation of GA lesions on images using a multimodal approach and a neural network system. This multimodal approach uses FAF images and one or both of IR (e.g., near-infrared (NIR)) and OCT images (e.g., frontal OCT images).

[0026] For example, near-infrared reflection imaging uses longer wavelengths than FAF to avoid medial occlusion, neurosensory layers, and maculaal luteinization. Therefore, GA lesions may appear brighter than non-atrophic areas. In such cases, NIR images can complement FAF images to facilitate the detection of foveal lesion boundaries, which may be more difficult with FAF alone due to lower image contrast / intensity in the fovea. In some cases, OCT frontal images, which are lateral images of the retinal and choroidal layers at a specified depth, can also be segmented in combination with FAF images, and possibly NIR images. Such a multimodal approach, in which images obtained from multiple imaging techniques are used as image input to an AI system for GA lesion segmentation (e.g., a neural network system), can facilitate the generation of improved segmentation results as described above, but can also be used to evaluate GA lesion expansion over time, i.e., over the long term.

[0027] Furthermore, this disclosure provides a system and method for automated segmentation of GA lesions using machine learning. In one embodiment, this disclosure provides a system and method for image segmentation using a multimodal approach and a neural network system. That is, for example, a neural network system can receive multimodal image inputs, i.e., image inputs including retinal images from two or more modalities (e.g., FAF images, NIR images, OCT images, etc.) can produce or generate GA lesion segmentation results that may be more accurate and beneficial compared to segmentation results from a single modality image input. This may be because retinal images from one modality may contain information about the depicted retina that may not be available from another modality.

[0028] Neural network systems may include, but are not limited to, convolutional neural network (CNN) systems and deep learning systems (e.g., U-Net deep learning neural network, Y-Net deep learning neural network, etc.). Such multimodal approaches that utilize neural network systems to segment retinal images facilitate accurate segmentation of GA lesions, as well as accurate and efficient long-term quantification of GA lesions, i.e., accurate and efficient quantification of GA lesion expansion over time.

[0029] For example, IR images, particularly NIR images, can provide a larger field of view than OCT frontal images. Furthermore, using IR images, more specifically NIR images, in combination with FAF images provides higher resolution and clarity for segmenting GA lesions. This higher resolution and clarity facilitates improved segmentation of GA lesions from retinal images, and consequently, improved feature extraction. For example, but not limited to, features such as the boundaries, shape, and texture of GA lesions can be more accurately identified. This improved feature extraction, in turn, enables improved overall long-term monitoring of GA lesions, which may be important for evaluating disease progression, treatment regimen effectiveness, or both.

[0030] For example, the correlation between actual GA lesion expansion from a baseline point to a later point in time (e.g., 6 months, 12 months, etc.) and GA lesion expansion estimated by the embodiments disclosed herein (i.e., using FAF images and both or either of IR images, particularly NIR and OCT images) can be higher than the correlation provided by some currently available methods. In some cases, the embodiments described herein allow for improved correlation with longer time intervals. For example, using both FAF images and IR images, particularly NIR images, the correlation provided for a baseline interval up to 12 months can be greater than the correlation provided for a baseline interval up to 6 months.

[0031] II. Definition This disclosure is not limited to these exemplary embodiments and uses, or the ways in which these exemplary embodiments and uses operate or are described herein. Furthermore, figures may be simplified or partial, and the dimensions of elements in the figures may be exaggerated or disproportionate.

[0032] Furthermore, wherever the terms “on,” “attached to,” “connected to,” “coupled to,” or similar terms are used herein, one element (e.g., a component, material, layer, substrate, etc.) can be “on,” “attached to,” “connected to,” or “coupled to” another element, regardless of whether one element is directly on top of another element, directly attached to another element, connected to another element, or coupled to another element, or whether one or more intervening elements exist between one element and the other. Furthermore, wherever a list of elements (e.g., elements a, b, c) is referenced, such reference is intended to include any one of the enumerated elements, any combination of fewer elements than all of the enumerated elements, and / or all combinations of the enumerated elements. The division of sections herein is merely for the convenience of consideration and does not limit any combination of elements described.

[0033] The term “subject” may refer to a subject in a clinical trial, a person undergoing treatment, a person receiving therapy, a person being monitored for remission or recovery, a person undergoing a preventive health analysis (e.g., due to their medical history), a person suffering from GA, or any other person or patient of interest. In various contexts, “subject” and “patient” may be used interchangeably herein.

[0034] Unless otherwise defined, scientific and technical terms used in connection with these instructions herein shall have meanings generally understood by those skilled in the art. Furthermore, unless otherwise required by context, singular terms shall include plural forms, and plural terms shall include singular forms. In general, nomenclature and techniques used in connection with chemistry, biochemistry, molecular biology, pharmacology, and toxicology are described herein, are well known and commonly used in the art.

[0035] As used herein, “substantially” means sufficient to function for the intended purpose. Thus, the term “substantially” allows for minor, slight variations from absolute or perfect conditions, dimensions, measurements, results, etc., which are expected by those skilled in the art but do not significantly affect the overall performance. When used in relation to numerical values, or parameters or characteristics that can be expressed numerically, “substantially” means within 10 percent.

[0036] As used herein, the term “about” when used with respect to a numerical value or a parameter or characteristic that can be expressed as a numerical value means within 10% of the numerical value. For example, “about 50” means a value in the range of 45 or more and 55 or less.

[0037] The term "plural" means two or more.

[0038] As used herein, the term “plural” may mean two, three, four, five, six, seven, eight, nine, ten or more.

[0039] As used herein, the term "set" means one or more items. For example, a set of items includes one or more items.

[0040] As used herein, the phrase “at least one of” means, when used with a list of items, that one or more different combinations of the enumerated items may be used, and only one of the items in the list may be required. An item can be a specific object, thing, step, action, process, or category. In other words, “at least one of” means that any combination or any number of items from the list may be used, but not all of the items in the list are required. For example, but not limited to, “at least one of item A, item B, or item C” means item A, item A and item B, item B, item A, item B, and item C, item B and item C, or item A and C. In some cases, “at least one of item A, item B, or item C” means, but not limited to, two of item A, one of item B, and ten of item C, four of item B and seven of item C, or several other suitable combinations.

[0041] As used herein, “model” may include one or more algorithms, one or more mathematical techniques, one or more machine learning algorithms, or a combination thereof.

[0042] As used herein, “machine learning” can include practices that use algorithms to analyze data, learn from it, and then make decisions or predictions about something in the world. Machine learning uses algorithms that can learn from data without relying on rule-based programming.

[0043] As used herein, “artificial neural network” or “neural network” (NN) may refer to a mathematical algorithm or computational model that mimics an interconnected group of artificial neurons that process information based on a connectivity-theoretic approach to computation. A neural network, sometimes called a neural net, can use one or more layers of nonlinear units to predict the output of an incoming input. Some neural networks include one or more hidden layers in addition to an output layer. The output of each hidden layer is used as input to the next layer in the network, i.e., the next hidden layer or output layer. Each layer of the network produces an output from an incoming input according to the current values ​​of each set of parameters. In various embodiments, a reference to “neural network” may refer to one or more neural networks.

[0044] Neural networks can process information in two ways: they are in training mode when they are being trained, and they are in inference (or prediction) mode when they actually perform what they have learned. Neural networks learn through a feedback process (e.g., backpropagation) that allows the network to adjust the weight coefficients of individual nodes in the intermediate hidden layers (correcting their behavior) so that the output matches the output of the training data. In other words, a neural network learns by being fed training data (learning examples) and eventually learns how to arrive at the correct output even when presented with a new range or set of inputs. A neural network can include, for example, at least one of the following types of neural networks, but are not limited to: feedforward neural networks (FNNs), recurrent neural networks (RNNs), modular neural networks (MNNs), convolutional neural networks (CNNs), residual neural networks (ResNets), ordinary differential equation neural networks (neural-ODEs), or other types of neural networks.

[0045] As used herein, “lesion” can include areas of organ or tissue that have been damaged through injury or disease. These areas can be continuous or discontinuous. For example, as used herein, a lesion can include multiple areas. A GA lesion is an area of ​​the retina that is suffering from chronic progressive degeneration. As used herein, a GA lesion can include one lesion (e.g., one continuous lesion area) or multiple lesions (e.g., a discontinuous lesion area consisting of multiple distinct lesions).

[0046] As used herein, “total lesion area” may refer to the area covered by the lesion (including the total area), whether the lesion is continuous or discontinuous.

[0047] As used herein, “long term” can mean a period of time. This period can be days, weeks, months, years, or any other measure of time.

[0048] As used herein, “encoder” may include a type of neural network that learns to efficiently encode data (e.g., one or more images) into a vector of parameters having multiple dimensions. The number of dimensions may be pre-selected.

[0049] As used herein, “decoder” may include a type of neural network that learns to efficiently decode a vector of parameters having several dimensions (e.g., several pre-selected dimensions) into output data (e.g., one or more images or image masks).

[0050] As used herein, “mask” may include an image of a type in which each pixel of the image has one of at least two different pre-selected potential values.

[0051] III. Segmentation of Geographic Atrophy (GA) Lesions Figure 1 is a block diagram of a lesion assessment system 100 according to various embodiments. The lesion assessment system 100 is used to assess geographic atrophy (GA) lesions in the retina of a subject. The lesion assessment system 100 includes a computing platform 102, data storage 104, and a display system 106. The computing platform 102 can take various forms. In one or more embodiments, the computing platform 102 includes a single computer (or computer system) or multiple computers communicating with each other. In other examples, the computing platform 102 takes the form of a cloud computing platform.

[0052] The data storage 104 and the display system 106 each communicate with the computing platform 102. In some examples, the data storage 104, the display system 106, or both may be considered part of the computing platform 102, or otherwise integrated. Thus, in some examples, the computing platform 102, the data storage 104, and the display system 106 may be separate components that communicate with each other, while in other examples, some combination of these components may be integrated together.

[0053] The lesion evaluation system 100 includes an image processor 108, which can be implemented using hardware, software, firmware, or a combination thereof. In one or more embodiments, the image processor 108 is implemented on a computing platform 102.

[0054] The image processor 108 receives an image input 109 for processing. In one or more embodiments, the image input 109 includes a set of fundus autofluorescence (FAF) images 110 and one or both of a set of infrared (IR) images 112 and a set of optical coherence tomography (OCT) images 124. In one or more embodiments, the set of IR images 112 is a set of near-infrared (NIR) images. In one or more embodiments, the image input 109 includes images produced by the same imaging device. For example, any combination of the set of FAF images 110, the set of IR images 112, and / or the set of OCT images 124 can be produced by the same imaging device. In one or more embodiments, any one of the set of FAF images 110, the set of IR images 112, and / or the set of OCT images 124 may be unaligned images. However, in other embodiments, any one of the set of FAF images 110, the set of IR images 112, and / or the set of OCT images 124 may be aligned images.

[0055] The image processor 108 processes the image input 109 (e.g., a set of FAF images 110, a set of IR images 112, and / or a set of OCT images 124) using the lesion quantification system 114 to generate a segmentation output 116 corresponding to the GA lesion. The lesion quantification system 114 may include any number or combination of neural networks. In one or more embodiments, the lesion quantification system 114 takes the form of a convolutional neural network (CNN) system comprising one or more neural networks. Each of these one or more neural networks may be a convolutional neural network itself. In some cases, the lesion quantification system 114 may be a deep learning neural network system. For example, the lesion quantification system 114 may be a Y-Net deep learning neural network, as will be described in more detail with reference to Figures 4A-4B. As another example, the lesion quantification system 114 may be a U-Net deep learning neural network, as will be described in more detail with reference to Figures 3A-3B.

[0056] In various embodiments, the lesion quantification system 114 may include a set of encoders 118 and a decoder 120. In one or more embodiments, the set of encoders 118 includes a single encoder, while in other embodiments, the set of encoders 118 includes multiple encoders. In various embodiments, each encoder in the set of encoders 118, as well as the decoder 120, may be implemented via a neural network, and the neural network may consist of one or more neural networks. In one or more embodiments, the decoder 120 and each encoder in the set of encoders 118 are implemented using a CNN. In one or more embodiments, the set of encoders 118 and the decoder 120 are implemented as a Y-Net (Y-shaped) neural network system or a U-Net (U-shaped) neural network system. For example, the lesion quantification system 114 may be a Y-Net neural network system having multiple encoders (e.g., two encoders) and fewer decoders than encoders (e.g., one decoder). As another example, the lesion quantification system 114 can be an equal number of encoders and decoders, for example, a U-Net neural network system having a single encoder and a single decoder.

[0057] The segmentation output 116 generated by the lesion quantification system 114 includes one or more segmentation masks. Each segmentation mask provides a pixel-level evaluation of a retinal region. For example, the segmentation mask in the segmentation output 116 may be a binary image in which each pixel has one of two values. In a specific example, the segmentation mask may be a black and white binary image in which white indicates regions identified as GA lesions. Non-limiting examples of segmentation masks indicating GA lesions are shown in Figures 6A and 6B.

[0058] In one or more embodiments, the lesion quantification system 114 is used to generate a preliminary probability map image in which each pixel has an intensity ranging from 0 to 1. Pixel intensities close to 1 are more likely to be GA lesions. The lesion quantification system 114 may include a thresholding module that applies a threshold to the preliminary probability map to generate a segmentation mask in the form of a binary probability map. For example, any pixel intensity in the preliminary probability map greater than or equal to a threshold (e.g., 0.5, 0.75, etc.) can be assigned an intensity of "1", and any pixel intensity in the preliminary probability map less than the threshold can be assigned an intensity of "0". Thus, the segmentation output 116 includes a binary segmentation mask that identifies regions identified as GA lesions.

[0059] In various embodiments, the image processor 108 (or another agent or module implemented within the computing platform 102) uses the segmentation output 116 to generate a quantitative assessment of GA lesions. For example, the image processor 108 can use the segmentation output 116 to extract a set of features 122 corresponding to GA lesions. The set of features 122 can be used to assess GA lesions (e.g., over time). For example, a set of FAF images 110, a set of IR images 112, and / or a set of OCT images 124 may include corresponding images at the same or substantially the same point in time (e.g., within the same time, same day, same 1-3 days, etc.). In one or more embodiments, each corresponding FAF image, as well as one or both of the IR and OCT images, may result in a corresponding segmentation mask in the segmentation output 116. The set of features 122 extracted from each of the different segmentation masks generated can provide or enable a long-term quantitative assessment of GA lesions over time.

[0060] The set of features 122 may include, but are not limited to, shape, texture, border or boundary map, number of lesions, total lesion area (i.e., area of ​​lesions when considered as composite or single lesion components), total lesion perimeter (i.e., perimeter of a single lesion component), ferret diameter of a single lesion component, excess rim intensity (e.g., the difference between the average intensity of the rim 0.5 mm around the lesion and the average intensity of the rim 0.5 to 1 mm around the lesion), roundness of the total lesion area, metrics indicating the subject's current or predicted / forecasted GA progression, or combinations thereof. These different types of features can be used to quantitatively evaluate GA lesions over time. For example, in the case of a segmentation mask within a segmentation output 116, the number of lesions may be the number of discontinuous regions or areas identified within the segmentation mask that form GA lesions. Total lesion area may be identified as the area or space occupied by one or more identified lesions, including the total area or space occupied by the lesions. Total lesion perimeter can be, for example, the perimeter of the general area or space occupied by one or more lesions. In other examples, total lesion perimeter can be the sum of the individual perimeters of one or more lesions. The metric indicates the subject's current or expected / predicted GA progression. In some cases, the metric may be calculated from a set of arbitrary features extracted from a segmentation mask, including, for example, shape, texture, border or boundary map, number of lesions, total lesion area (i.e., the area of ​​lesions if considered as composite or single lesion components), total lesion perimeter (i.e., the perimeter of a single lesion component), ferret diameter of a single lesion component, excess rim intensity (e.g., the difference between the average intensity of the rim 0.5 mm around the lesion and the average intensity of the rim 0.5-1 mm around the lesion), roundness of the total lesion area, or a combination thereof.

[0061] In one or more embodiments, the first segmentation mask in the segmentation output 116 may correspond to the baseline time, the second segmentation mask in the segmentation output 116 may correspond to 6 months after baseline, and the third segmentation mask in the segmentation output 116 may correspond to 12 months after baseline. A set of features 122 (e.g., shape, texture, number of lesions, total lesion area, total lesion perimeter, ferret diameter, excess rim intensity, circularity of lesion area, etc.) can be identified for each of the first, second, and third segmentation masks. Thus, the set of features 122 can be used to quantitatively evaluate GA lesions over a period between baseline and 12 months.

[0062] In some embodiments, the set of features 122 includes features corresponding to a certain time range. With respect to the above examples illustrating the first, second, and third segmentation masks, the set of features 122 may include the expansion or growth rate of GA lesions over a 6-month period, over a 12-month period, or both. The rate of change in lesion area, i.e., the expansion rate, can be calculated using one or more other identified features. For example, the expansion rate over 6 months can be calculated based on the difference between the total lesion area extracted for the first segmentation mask and the total lesion area extracted for the second segmentation mask. The expansion rate over a 12-month period can be calculated based on the difference between the total lesion area extracted for the first segmentation mask and the total lesion area extracted for the third segmentation mask. In this way, the set of features 122 may include one or more features that are calculated based on one or more other features within the set of features 122.

[0063] The segmentation output 116 may include a number of segmentation masks necessary to evaluate GA lesions over a selected period (e.g., 3 months, 6 months, 12 months, 18 months, etc.). Furthermore, the segmentation output 116 may include a desired number of segmentation masks to evaluate GA lesions at desired time intervals within the selected period. The desired time intervals may be constant or different. In one or more embodiments, the segmentation output 116 includes a 10-day segmentation mask over a 12-month period.

[0064] Figure 2 is a flowchart of process 200 for evaluating geographic atrophic lesions according to various embodiments. In various embodiments, process 200 is implemented using the lesion evaluation system 100 described in Figure 1.

[0065] Step 202 involves receiving a set of fundus autofluorescence (FAF) images of the retina having one or more geographic atrophy (GA) lesions.

[0066] Step 204 includes receiving either or both a set of infrared (IR) images of the retina or a set of optical coherence tomography (OCT) images of the retina.

[0067] Step 206 includes generating a GA lesion segmentation mask containing one or more GA lesion segments corresponding to one or more GA lesions in the retina, using a set of FAF images and one or both of the set of IR images or the set of OCT images. In some embodiments, generating includes generating the GA lesion segmentation mask using a neural network, the neural network including a U-Net deep learning neural network having an encoder and a decoder. In such cases, process 200 may further include generating an encoded image input by concatenating the set of FAF images and one or both of the set of IR images or the set of OCT images using the encoder. In some cases, generating the GA lesion segmentation mask includes decoding the encoded image input using the decoder of the U-Net deep learning neural network.

[0068] In some embodiments, generating includes generating a GA lesion segmentation mask using a neural network, the neural network including a Y-Net deep learning neural network having a first encoder, a second encoder or a third encoder (or both), and a decoder. In such cases, process 200 further includes generating an encoded FAF image input using a set of FAF images via the first encoder, generating an encoded IR image input from a set of IR images via the second encoder, or generating an encoded OCT image input from a set of OCT images via the third encoder (or both). Furthermore, in some cases, process 200 includes generating an encoded image input by concatenating the encoded FAF image input with one or both of the encoded IR image input or the encoded OCT image input, and generating a GA lesion segmentation mask by decoding the input encoded image using the decoder of the Y-Net deep learning neural network.

[0069] Some embodiments of process 200 further include the processor extracting features of one or more GA lesions in the retina from a GA lesion segmentation mask. Furthermore, process 200 may also include the processor generating recommendations for treating one or more GA lesions based on the extracted features. In some embodiments, the extracted features include several one or more GA lesions.

[0070] Some embodiments of process 200 further include the processor combining one or more GA lesion segments into a single lesion component, the extracted features including one or more of the area, perimeter, ferret diameter, or excess rim intensity of the single lesion component.

[0071] Figure 3A shows a process 300 that utilizes a U-Net deep learning neural network to evaluate geographic atrophy lesions according to various embodiments. In some embodiments, process 300 can use a multimodal approach in which images of the retina with one or more GA lesions are obtained by using multiple different imaging techniques for segmentation via a U-Net deep learning neural network (also abbreviated as "U-Net"). In some cases, U-Net can receive a set of FAF images 310a of the retina with one or more GA lesions, for example from cSLO. U-Net can also receive a set of IR images 310b of the retina from an IR imaging device (e.g., NIR images) and / or a set of OCT images 310c of the retina from an OCT imaging device (e.g., volumetric OCT images).

[0072] In some embodiments, one or more of the received multimodal image inputs, namely one or more of the set of FAF images 310a, the set of IR images 310b, and the set of OCT images 310c, may be preprocessed 320a, 320b, 320c before being integrated as a multichannel image input 330. For example, the image inputs may be resized, or their intensity may be normalized (e.g., to a scale of 0 to 1). In the case of the set of OCT images 310c, preprocessing 320c may include applying histogram matching to the volume OCT image 310c and flattening each B scan along the Bruch membrane to generate a frontal OCT image. In some cases, the frontal OCT images can be combined to generate a multichannel frontal OCT image input. For example, three frontal images or maps may be averaged over the entire depth on the Bruch membrane, and the sub-Bruch membrane depth may be combined with a 3-channel frontal OCT image input. In some cases, a set of pre-processed multimodal image inputs may be integrated (e.g., concatenated) as a multi-channel image input 330 that can then be encoded by the U-Net encoder 340.

[0073] An exemplary architecture of a U-Net that can be used to segment retinal images as discussed in this disclosure is presented below with reference to a multimodal image input including a set of FAF images and a set of NIR images. However, it should be understood that the exemplary architecture and related descriptions are for illustrative purposes only, and this disclosure is applicable to any U-Net architecture applied to a multimodal image input including a set of retinal images of different numbers and types. In some embodiments, the U-Net architecture can be designed to predict and classify each pixel in an image (e.g., a multi-channel image input 330), which allows for more accurate segmentation with fewer training images. In some cases, the U-Net may include a collision encoder E340 and an expansion decoder D350. In this architecture, a set of preprocessed images (e.g., a set of FAF images 310a, a set of IR images 310b, and a set of OCT images 310c after the respective preprocessing steps 320a, 320b, and 320c) can be combined or concatenated to generate a multi-channel input 330, which is then encoded by the collision encoder E340. In some cases, the encoded image may be passed to the extended decoder D350 to generate the segmentation mask 360. Referring to a multimodal image input including FAF image 310a and NIR image 310b as an illustrative example, the process can be represented as follows: (Z,S=E(concat(FAF,NIR)); P=D(Z,S), Here, "FAF" represents an FAF image, and "NIR" represents an NIR image.

[0074] An exemplary embodiment of the U-Net architecture is shown in Figure 3B, where the erosion encoder E340 alternates between convolutional blocks and downsampling operations, allowing for a total of six downsamples. Each convolutional block contains two 3x3 convolutions. After each convolution, group normalization (GroupNorm) and normalized linear unit (ReLU) activation can be performed. Downsampling can be performed by a 2x2 maxpool operation. Z can be the final encoded image, and S can be a set of partial results from the erosion encoder E340 prior to each downsampling step. In the decoder D350, upsampling steps and convolutional blocks are performed alternately, allowing for a total of six upsamples. In some cases, the convolutional blocks can be the same as or similar to the convolutional blocks in the encoder, but the upsampling can be performed by a 2x2 transposed convolution. After each upsampling, partial results of the same size from S can be copied in a skip connection and concatenated into the result. After the final convolution block, a 1x1 convolution is performed, followed by sigma activation to obtain a probability P between 0 and 1, from which a segmentation mask 360 can be generated. For example, image pixels with a probability P above a threshold are assigned a value of "1" and identified as GA lesions, while pixels with a probability P below a threshold are assigned a value of "0" indicating the absence of GA lesions, resulting in a binary segmentation map as shown in Figures 6A and 6B.

[0075] In some embodiments, due to the encoder-decoder structure of the U-Net shown in Figure 3B and the number and size of convolutional filters used, the U-Net deep learning neural network can learn image context information. In some cases, upsampling in decoder D350 can help propagate image context information to higher resolution channels. Image context information, or features, can be thought of as miniimages, which are small two-dimensional arrays of values. Features can match common aspects of an image. For example, in the case of an image of the letter "X", features consisting of diagonals and intersections can capture most or all of the important properties of X. In some cases, these features can match up to any image of X. Thus, features are essential values ​​of miniimages that capture some context information of an image. The U-Net architecture generates a large number of such features, enabling the U-Net deep learning neural network to learn the relationship between images and their corresponding labels / masks.

[0076] Returning to Figure 3A, in some embodiments, features 370 of one or more GA lesions in the retina can be extracted from the segmentation mask 360. For example, the number of one or more GA lesions may be determined from the segmentation mask 360. Other features 370 that can be extracted or determined from the segmentation mask 360 include, but are not limited to, the size, shape, area, perimeter, ferret diameter, roundness, and excess rim intensity of the GA lesions. In some cases, one or more GA lesions may be fragmented, in which case the features 370 above may refer to parameters of a single lesion component consisting of one or more GA lesions. That is, these parameters (i.e., size, shape, area, perimeter, ferret diameter, roundness, excess rim intensity, etc.) can be extracted or determined for a single lesion component combining one or more GA lesions. In some cases, recommendations for treating one or more GA lesions in the retina may be determined or generated based on the extracted features 370. For example, an ophthalmologist can diagnose a GA lesion (e.g., the stage or severity of the GA lesion) and prescribe treatment based on the extracted characteristics 370.

[0077] Figure 4A shows a process 400 that utilizes a Y-Net deep learning neural network to evaluate geographic atrophy lesions according to various embodiments. In some embodiments, process 400 can use a multimodal approach in which images of the retina with one or more GA lesions are obtained by using multiple different imaging techniques for segmentation via a Y-Net deep learning neural network (also abbreviated as "Y-Net"). In some cases, Y-Net can receive a set of FAF images 410a of the retina with one or more GA lesions, for example from cSLO. Y-Net can also receive a set of IR images 410b of the retina from an IR imaging device (e.g., NIR images) and / or a set of OCT images 410c of the retina from an OCT imaging device (e.g., volumetric OCT images).

[0078] In some embodiments, one or more of the received multimodal image inputs, namely one or more of the set of FAF images 410a, the set of IR images 410b, and the set of OCT images 410c, may be preprocessed 420a, 420b, 420c before being encoded by the respective Y-Net encoders 430a, 430b, 430c. For example, the image inputs may be resized, or their intensity may be normalized (e.g., to a scale of 0 to 1). In the case of the set of OCT images 410c, preprocessing 420c may include applying histogram matching to the volume OCT image 410c and flattening each B scan along Bruch's membrane to generate a frontal OCT image. In some cases, the frontal OCT images can be combined to generate a multichannel frontal OCT image input. For example, three frontal images or maps may be averaged over the entire depth on the Bruch membrane, and the sub-Bruch membrane depth may be combined with three-channel frontal OCT image inputs. In some cases, as described above, a set of preprocessed multimodal image inputs (e.g., a set of FAF images 410a, a set of IR images 410b, and a set of frontal OCT images 410c) may be encoded by the respective Y-Net encoders 430a, 430b, and 430c.

[0079] In some cases, the Y-Net may have as many encoders as there are modalities that are the source of the retinal images. For example, if the multimodal image input to the Y-Net includes a pair of retinal image sets (e.g., a pair containing a set of FAF images 310a and a set of IR images 310b (e.g., not a set of OCT images 310c), or a pair containing a set of FAF images 310a and a set of OCT images 310c (e.g., not a set of IR images 310b)), the Y-Net may be configured to have two encoders, each encoder configured to encode one of the pair of image sets. As another example, if the multimodal image input to the Y-Net includes three retinal image sets (e.g., a set of FAF images 310a, a set of IR images 310b, and a set of frontal OCT images 310c), the Y-Net may be configured to have three encoders (e.g., encoders 430a, 430b, and 430c), each encoder configured to encode one of the three image sets.

[0080] An exemplary architecture of Y-Net that can be used to segment retinal images as discussed in this disclosure is presented below with reference to a multimodal image input including a set of FAF images and a set of NIR images. However, it should be understood that the exemplary architecture and related descriptions are for illustrative purposes only, and this disclosure is applicable to any Y-Net architecture applied to a multimodal image input including a different number and type of retinal image sets. In some embodiments, the Y-Net may include a decoder D440 and as many encoder branches as there are modalities in the image input (e.g., two encoders E1430a and E2430b). In this architecture, the preprocessed sets of images (e.g., a set of FAF images 410a, a set of IR images 410b, and a set of OCT images 410c after the respective preprocessing steps 420a, 420b, and 420c) can be encoded separately by their respective encoders. For example, referring to Figure 4A, the set of FAF images 410a after preprocessing 420a and the set of IR images 410b after preprocessing 420b may be encoded separately by encoder E1430a and encoder E2430b, respectively, before being combined or concatenated before decoding by the Y-Net decoder D440, from which a segmentation mask 450 may be generated. Referring to a multimodal image input including FAF images 410a and NIR images 410b as an illustrative example, the process can be represented as follows: (Z1,S1=E1(FAF); (Z2,S2=E2(NIR); P=D(concat(Z1,Z2),S1).

[0081] In some cases, encoders E1430a and E2430b may have the same or similar architecture as encoder 340 in the U-Net shown in Figures 3A and 3B, except that the image inputs of encoders E1430a and E2430b are single-channel image inputs, i.e., a set of pre-processed FAF images 410 and a set of NIR images 420a, respectively, and the image input of encoder 340 is a multi-channel image input 330. Furthermore, decoder D440 may have the same or similar architecture as decoder D350, except that the input of the former may have the same number of channels as the number of image input modalities. For example, if the multimodal image input includes FAF images 410a and NIR images 410b, as shown in the exemplary figures above, the input of decoder D440 may have twice the number of channels as encoder 340. Furthermore, in some cases, a partial result S1 from the first encoder E1 (e.g., FAF encoder E1) was used in decoder D440, but a partial result from the second encoder E2 (e.g., NIR encoder E2) was not used. Then, a probability P having a value between 0 and 1 can be calculated from decoder D440, from which a segmentation mask 450 can be generated as described above with reference to Figure 3A. That is, image pixels with a probability P above the threshold are assigned a value of "1" and identified as GA lesions, and pixels with a probability P below the threshold are assigned a value of "0" indicating that there are no GA lesions, resulting in a binary segmentation map as shown in Figures 6A-6B.

[0082] As described above regarding the encoder-decoder structure of U-Net and Y-Net in Figure 4B, and the number and size of convolutional filters used, the Y-Net deep learning neural network can learn image context information. In some cases, upsampling in decoder D440 can help propagate image context information to higher resolution channels. Image context information, or features, can be thought of as miniimages, which are small two-dimensional arrays of values. Features can match common aspects of an image. For example, in the case of an image of the letter "X", features consisting of diagonals and intersections can capture most or all of the important properties of X. In some cases, these features can match up to any image of X. Thus, features are essential values ​​of miniimages that capture some context information of an image. The Y-Net architecture generates a large number of such features, enabling the Y-Net deep learning neural network to learn the relationship between images and their corresponding labels / masks.

[0083] In some embodiments, as described above with reference to Figure 3A, features 460 of one or more GA lesions in the retina can also be extracted from the segmentation mask 450. For example, the number, size, shape, area, perimeter, ferret diameter, roundness, excess rim strength, etc., of one or more GA lesions can be determined from the segmentation mask 450. If, as described above, one or more GA lesions are fragmented, the features 460 above may refer to parameters of a single lesion component consisting of one or more GA lesions. That is, these parameters (i.e., size, shape, area, perimeter, ferret diameter, roundness, excess rim strength, etc.) can be extracted or determined for a single lesion component combining one or more GA lesions. Furthermore, in some cases, recommendations for treating one or more GA lesions in the retina may be determined or generated based on the extracted features 460. For example, an ophthalmologist can diagnose one or more GA lesions and prescribe treatment based on the extracted features 460 (e.g., based on the size of the GA lesions).

[0084] In some embodiments, the U-Net deep learning neural network in Figures 3A-3B and / or the Y-Net deep learning neural network in Figures 4A-4B may be trained to generate segmentation masks 360 and 450, respectively, as follows. In some cases, the U-Net and Y-Net architectures may be initialized with random weights. The loss function to be minimized can be a dice loss defined as 1 minus the dice coefficient. The Adam optimizer, an algorithm for first-order gradient-based optimization of a stochastic objective function, can be used for training. The initial learning rate is 1e -3It is set to (i.e., 0.001) and can be multiplied by 0.1 every 30 epochs for 100 epochs without premature termination. The hyperparameters can then be tuned and selected using the validation dataset. Once the hyperparameters are selected, in some embodiments, U-Net and Y-Net can be used to predict the GA lesion segmentation mask on the test set. To evaluate long-term outcomes, GA lesion expansion is measured in the test set as the absolute change (mm²) of GA lesion area over time, for example, from baseline to 6 months and from baseline to 12 months. 2 It can be calculated as follows. Furthermore, the GA lesion expansion of the predicted segmentation of U-Net and / or Y-Net can be compared to that of the evaluator annotation. The Dice score can also be calculated and reported as a performance metric for CNN segmentation. In addition, the Pearson correlation coefficient (r) can be used to assess long-term performance.

[0085] In some embodiments, during the training of a neural network, a modified version of the ground truth mask can be used to weight the edges of lesions more than the interiors. For example, the modified mask may be defined as follows: TIFF0007853292000001.tif21170 Here, orig_mask represents the original ground truth mask with values ​​of 0 and 1, and edt represents the Euclidean distance transformation. In some cases, all operations except edt and max may be element-wise operations, with max returning the maximum value of the input array. The result was a mask of 0 if there were no lesions, and at least 0.5 if there were lesions. In some cases, higher values ​​(up to 1) may be recorded near the edges of lesions.

[0086] In some cases, the following dice function can be used: TIFF0007853292000002.tif12170 Here, ε is a parameter used to ensure smoothness when the mask is empty (e.g., ε=1). The gradient of the dice function with this modification can help the neural network (U-Net or Y-Net) output the correct probability value P between 0 and 1 (i.e., the correct segmentation mask). In other words, the neural network can, with the help of the gradient of the dice function, compute an increased probability value P (e.g., down to a value of 1) for regions or pixels on the input image that have GA lesions, and a decreased probability P (e.g., down to a value of 0) for regions or pixels that do not have GA lesions. In some cases, the neural network can place more emphasis on the boundaries of the GA lesions rather than the interior, which can be appropriate because the neural network may find it easier to properly identify the interior of the GA lesions, and therefore the neural network may not spend as much time learning to do so.

[0087] In some embodiments, the original mask can be used during verification and testing, and the prediction or probability may have a value of 1 if the predicted probability is greater than a threshold, and a value of 0 otherwise. For example, if the prediction or probability is greater than or equal to the threshold of 0.5, the probability P may be assigned a value of 1 indicating the presence of a GA lesion, and if the prediction or probability is less than the threshold of 0.5, the probability P may be assigned a value of 0 indicating the absence of a GA lesion.

[0088] IV. Exemplary applications of the systems and methods disclosed herein Figure 5 illustrates an exemplary workflow for a retrospective study using Y-Net and U-Net deep learning neural networks to segment FAF and NIR images, relating to various embodiments. This study was conducted using FAF and NIR imaging data from study eyes of patients enrolled in an observational clinical study titled "Study of Disease Progression in Participants with Geographic Atrophy Secondary to Age-Related Macular Degeneration" (a so-called Proxima A and Proxima B natural history study of GA patients). Proxima A patients had bilateral GA without choroidal neovascularization (CNV) at baseline, while Proxima B patients had unilateral GA with or without CNV in the other eye at baseline.

[0089] FAF images at screening and follow-up visits were graded by two junior evaluators. If the junior evaluators did not agree, a senior grader also graded the scans. The results shown in Figures 6–11 use only scans graded by both junior evaluators (one patient was not included due to incomplete grading from the second evaluator). FAF images from the Proxima B study were graded by a senior evaluator. As shown in Figure 5, 940 FAF and NIR image pairs were obtained from 194 patients in the Proxima B study, lesions were annotated on the FAF images by human evaluators, and the data was split at the patient level into a training set (images from 155 patients) with a total of 748 visits and a validation set (images from 39 patients) with a total of 192 visits. 90 FAF and NIR image pairs from 90 patients were used in the study. Visits from patients without both FAF and NIR images were excluded. The U-Net and Y-Net neural networks were trained using the training set as described above and validated using the validation set.

[0090] In the studies of both sides, GA diagnosis and lesion area measurement were based on FAF imaging when the minimum lesion diameter was used as the cutoff for recognizing a GA lesion as one. In Proxima A, the study eye had a well-defined area of GA secondary to AMD, without evidence of previous CNV or active CNV, and the total GA lesion size was more than 2.54 mm 2 but less than 17.78 mm 2 and was completely within the FAF imaging field (field of view 2 - 30°, image centered on the fovea) and had a striated or diffuse hyperautofluorescence pattern around the lesion, the GA lesion was recognized as one. If the GA lesion was multifocal, the GA lesion was recognized as one when at least one nodular lesion had an area of 1.27 mm 2 or more. In Proxima B, the GA lesion was recognized as one in the study eyes of patients with GA and without CNV, and in fellow eyes with or without GA and with a total lesion size of CNV from 1.27 mm 2 to 17.78 mm 2 . In patients with unilateral GA and no CNV in the study eye, the GA lesion was recognized as one when its total lesion size was from 0.3 mm 2 to 17.78 mm 2 , or when it was multifocal and at least one nodular lesion had an area of 0.3 mm 2 or more.

[0091] The GA lesion size was obtained from GA lesions annotated by human assessors (i.e., considered to be ground truth) using RegionFinder software. The minimum size of an individual lesion was 0.05 mm 2The settings were adjusted to a diameter of approximately 175 μm. FAF and NIR images from the test set from the Proxima A study were segmented using U-Net and Y-Net as described above with reference to Figures 3 and 4, respectively, and the outputs from these neural networks were then compared to ground truth (i.e., annotations from two evaluators G1 and G2). Corresponding infrared images were used to identify small atrophy areas and atrophy depictions around the fovea. Images and masks were resized to 768 × 768 pixels and no normalization was performed.

[0092] Figures 6–11 show various measurements extracted from GA lesion segmentation masks generated by human evaluators and U-Net / Y-Net neural networks, obtained from FAF and NIR images of the Proxima A study. Figures 6A–6B show exemplary segmentation results of FAF and NIR images of retinas with GA lesions using two human evaluators G1 and G2, as well as Y-Net and U-Net deep learning neural networks, according to various embodiments. In particular, in both Figures 6A and 6B, the first row shows the FAF and NIR images of the retina at screening, along with the GA lesion segmentation results performed by each human evaluator and neural network, as well as dice scores (shown in more detail in Figure 7) measuring the similarity of the segmentation performed by the evaluators and neural networks. The second row shows the corresponding results for FAF and NIR images of the same retina taken 12 months later.

[0093] Figure 6A shows good agreement between human evaluators and neural network segmentation results, confirmed by fairly high Dice scores, while Figure 6B shows poorer agreement, again confirmed by lower Dice scores. The worse agreement shown in Figure 6B may be related to foveal evaluation, where the U-Net and Y-Net neural networks identify the fovea as lesions. The Y-Net neural network appears to perform better than the U-Net neural network in avoiding foveal segmentation. However, the differences with Y-Net and any foveal-specific advantages may be minor. Poor agreement can also be attributed to the neural network misinterpreting shadows as lesions or to poor FAF image quality. Note that while there is a high correlation in lesion area, the contours are not identical.

[0094] Figure 7 shows exemplary Dice similarity coefficient scores for measuring the similarity between Y-Net deep learning neural networks, U-Net deep learning neural networks, and segmentations performed by human evaluators, according to various embodiments. In particular, Figure 7 shows a table containing Dice scores that measure the similarity between GA lesion segmentations performed by the first assessor G1 and Y-Net ("G1-YNet"), the first assessor G1 and Y-Net ("G1-UNet"), the second assessor G2 and Y-Net ("G1-YNet"), the second assessor G2 and U-Net ("G1-UNet"), and the first assessor G1 and second assessor G2 ("G1-G2") at the time of initial screening ("SCR"), 6 months after screening ("M6"), 12 months after screening ("M12"), 18 months after screening ("M18"), 24 months after screening ("M24"), ("ET"), and combinations thereof ("All"). Furthermore, Figure 7 shows a swarmplot containing the aforementioned Dice scores at the time of screening. The results shown in the figure indicate that the agreement between GA lesion segmentation performed by the neural network and that performed by human evaluators was similar to the agreement between GA lesion segmentation performed by two evaluators.

[0095] Figures 8A–8D illustrate illustrative comparisons of screening GA lesion area sizes measured by a Y-Net deep learning neural network, a U-Net deep learning neural network, and a human evaluator, according to various embodiments. In some embodiments, to evaluate the performance of the neural network for individual GA lesion areas, the correspondence of GA lesion areas within segmentation masks generated by the neural network was compared to the average GA lesion area of ​​segmentation masks annotated by the evaluator. Figures 8A and 8B both show good agreement between screening GA lesion areas in segmentation masks generated by Y-Net and U-Net, respectively, and the average GA lesion area of ​​segmentation masks annotated by the evaluator ("average evaluator"). Figure 8C also shows good agreement between screening GA lesion areas in segmentation masks annotated by two evaluators. Figure 8D shows a Bland-Altman plot, i.e., a difference plot showing the difference between the GA lesion area within the segmentation mask annotated by the raters and generated by the neural network, and the mean GA lesion area (y-axis) as a function of the mean (x-axis). This figure also shows a polynomial line with a smoothness of 2 to show the general trend. However, it should be noted that these plots may have inherent statistical bias because the rater mean is used as the baseline, and the raters were a priori closer to the baseline than the neural network.

[0096] Figures 9A–9E show exemplary comparisons of 12-month lesion area sizes measured by the Y-Net deep learning neural network, the U-Net deep learning neural network, and human evaluators, according to various embodiments. In some embodiments, the ability of a neural network to measure GA lesion area change can be tested by measuring GA lesion area in the same eye or retina at different points in time. Figures 9A–9C show the change in annular GA lesion area (i.e., long-term area change) from screening time to 12 months later between segmentation performed by Y-Net and the "average evaluator" (Figure 9A), between segmentation performed by U-Net and the "average evaluator" (Figure 9B), and between segmentation performed by two evaluators (Figure 9C). Figure 9D shows a Bland-Altman plot annotated by the evaluators, showing the difference between the long-term change in GA lesion area within the segmentation mask generated by the neural network and the mean (y-axis) of the GA lesion area change as a function of the mean (x-axis). This figure also shows a polynomial line with a smoothness of 2 to illustrate the general trend. Figures 9A–9D show that the neural network-rater comparison results for long-term GA lesion area are poorer compared to cross-sectional results (such as those shown in Figures 8A–8D). This poor performance for long-term measurements may be a statistical artifact resulting from potentially small mean change magnitudes. Furthermore, the neural network may not have the opportunity to process each image independently compared to previous visits. Comparisons between screening and 6 months were also found to be even worse in general, potentially due to noise amplification and algorithmic errors, as well as small sample size. However, the time-series comparison of measured area changes shown in Figure 9E shows similar mean changes and CV between the network and raters, suggesting that the neural network performs well over time in endpoint measurements, although the correlation is not high at the individual patient level or short time windows.

[0097] V. Artificial Neural Networks Figure 10 shows an exemplary neural network that can be used to implement computer-based models according to various embodiments of the present disclosure. For example, the neural network 1000 can be used to implement the lesion quantification system 114 of the lesion assessment system 100. As shown in the figure, the artificial neural network 1000 includes three layers: an input layer 1002, a hidden layer 1004, and an output layer 1006. Each of layers 1002, 1004, and 1006 may include one or more nodes. For example, the input layer 1002 includes nodes 1008-1014, the hidden layer 1004 includes nodes 1016-1018, and the output layer 1006 includes node 1022. In this example, each node in the hierarchy is connected to all nodes in the adjacent hierarchy. For example, node 1008 in the input layer 1002 is connected to both nodes 1016 and 1018 in the hidden layer 1004. Similarly, node 1016 of the hidden layer is connected to all nodes 1008-1014 of the input layer 1002 and node 1022 of the output layer 1006. Although only one hidden layer is shown for the artificial neural network 1000, it is conceivable that the artificial neural network 1000 used to implement the lesion quantification system 114 of the lesion evaluation system 100 could include many hidden layers as needed or desired.

[0098] In this example, the artificial neural network 1000 receives a set of input values ​​and generates output values. Each node in the input layer 1002 can correspond to a distinct input value. For example, if the artificial neural network 1000 is used to implement the lesion quantification system 114 of the lesion evaluation system 100, each node in the input layer 1002 can correspond to a distinct attribute of a set of FAF images 110, a set of IR images 112, or a set of OCT images 124.

[0099] In some embodiments, each of the nodes 1016-1018 in the hidden layer 1004 may generate a representation, which may include a mathematical computation (or algorithm) that generates a value based on input values ​​received from nodes 1008-1014. The mathematical computation may include assigning different weights to each of the data values ​​received from nodes 1008-1014. Nodes 1016 and 1018 may include different algorithms and / or different weights assigned to the data variables from nodes 1008-1014 so that each of nodes 1016-1018 can generate different values ​​based on the same input values ​​received from nodes 1008-1014. In some embodiments, the weights initially assigned to each feature (or input value) of nodes 1016-1018 may be generated randomly (e.g., using a computer randomizer). The values ​​generated by nodes 1016 and 1018 can be used by node 1022 in the output layer 1006 to generate the output value of the artificial neural network 1000. If the artificial neural network 1000 is used to implement the lesion quantification system 114 of the lesion evaluation system 100, the output values ​​generated by the artificial neural network 1000 may include a segmentation output 116.

[0100] The artificial neural network 1000 can be trained by using training data. For example, the training data here may be a set of FAF images, a set of IR images, or a set of OCT images. By providing training data to the artificial neural network 1000, nodes 1016-1018 in the hidden layer 1004 can be trained (tuned) so that the output layer 1006 produces the optimal output based on the training data. By sequentially providing different training datasets and penalizing the artificial neural network 1000 when its output is incorrect (e.g., when it produces a segmentation mask containing an incorrect GA lesion segment), the artificial neural network 1000 (specifically, the representation of the nodes in the hidden layer 1004) can be trained (tuned) to improve its performance in data classification. Tuning the artificial neural network 1000 may include adjusting the weights associated with each node in the hidden layer 1004.

[0101] While the above description relates to artificial neural networks as an example of machine learning, it will be understood that other types of machine learning methods may also be suitable for implementing various aspects of this disclosure. For example, machine learning can be implemented using support vector machines (SVMs). SVMs are a set of related supervised learning methods used for classification and regression. An SVM training algorithm, which can be a non-stochastic binary linear classifier, can build a model that predicts whether a new example will fall into one category or another. Another example is the implementation of machine learning using Bayesian networks. A Bayesian network is an acyclic stochastic graphical model that represents a set of random variables and their conditional independence by a directed acyclic graph (DAG). A Bayesian network can present a stochastic relationship between one variable and another. Another example is a machine learning engine that performs a machine learning process using a decision tree learning model. In some cases, the decision tree learning model may include classification tree models and regression tree models. In some embodiments, the machine learning engine uses a gradient boosting machine (GBM) model (e.g., XGBoost) as the regression tree model. Other machine learning techniques can be used to implement a machine learning engine, for example, via a random forest or a deep neural network. Other types of machine learning algorithms are not described in detail herein for simplicity, and it should be understood that this disclosure is not limited to any particular type of machine learning.

[0102] VI. Computer-Implemented Systems Figure 11 is a block diagram of a computer system 1100 suitable for implementing various methods and apparatus described herein, such as the lesion assessment system 100, computing platform 102, data storage 104, display system 106, image processor 108, etc. In various implementations, the devices capable of performing the steps may include network communication devices (e.g., mobile phones, laptops, personal computers, tablets, etc.), network computing devices (e.g., network servers, computer processors, developer workstations, electronic communication interfaces, etc.), or other suitable devices. Therefore, it should be understood that the devices capable of implementing the servers, systems, and modules described above, as well as the various method steps of method 200 described above, can be implemented as computer system 1100 in the following ways.

[0103] According to various embodiments of this disclosure, a computer system 1100, such as a network server, workstation, computing device, or communication device, includes a bus component 1102 or other communication mechanism for communicating information, the bus component interconnecting subsystems and components such as computer processing components 1104 (e.g., processor, microcontroller, digital signal processor (DSP), etc.), system memory components 1106 (e.g., RAM), static storage components 1108 (e.g., ROM), disk drive components 1110 (e.g., magnetic or optical), network interface components 1112 (e.g., modem or Ethernet card), display components 1114 (e.g., cathode ray tube (CRT) or liquid crystal display (LCD)), input components 1116 (e.g., keyboard), cursor control components 1118 (e.g., mouse or trackball), and image capture components 1120 (e.g., analog or digital camera). In one implementation, the disk drive component 1110 may include a database having one or more disk drive components.

[0104] According to embodiments of the present disclosure, the computer system 1100 performs specific operations by a processor 1104 that executes one or more sequences of one or more instructions contained in a system memory component 1106. Such instructions may be read into the system memory component 1106 from another computer-readable medium, such as a static storage component 1108 or a disk drive component 1110. In other embodiments, hardwired circuits may be used instead of (or in combination with) software instructions to implement the present disclosure. In some embodiments, various components, such as a lesion assessment system 100, a computing platform 102, data storage 104, a display system 106, and an image processor 108, may be in the form of software instructions that can be executed by the processor 1104 to automatically perform context-appropriate tasks on behalf of the user.

[0105] The logic can be encoded into a computer-readable medium that can refer to any medium involved in providing instructions to the processor 1104 for execution. Such a medium can take many forms, including but not limited to non-volatile and volatile media. In one embodiment, the computer-readable medium is non-transient. In various implementations, the non-volatile medium includes optical or magnetic disks, such as disk drive component 1110, and the volatile medium includes dynamic memory, such as system memory component 1106. In one embodiment, data and information related to execution instructions may be transmitted to the computer system 1100 via a transmitting medium, such as in the form of sound waves or optical waves, including those generated during radio and infrared data communications. In various implementations, the transmitting medium may include coaxial cables, copper wires, and optical fibers, including wires with a bus 1102.

[0106] Some common forms of computer-readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tapes, any other magnetic media, CD-ROMs, any other optical media, punch cards, paper tapes, any other physical media having a pattern of holes, RAM, PROMs, EPROMs, FLASH-EPROMs, any other memory chips or cartridges, carriers, or any other media adapted for computer reading. These computer-readable media may also be used to store programming code for the aforementioned lesion assessment system 100, computing platform 102, data storage 104, display system 106, image processor 108, etc.

[0107] In various embodiments of the present disclosure, the execution of instruction sequences for implementing the present disclosure may be performed by a computer system 1100. In various other embodiments of the present disclosure, multiple computer systems 1100 connected by a communication link 1122 (e.g., a communication network such as a LAN, WLAN, PTSN, and / or various other wired or wireless networks including telecommunications, mobile, and cell phone networks) may cooperate with each other to execute instruction sequences for implementing the present disclosure.

[0108] The computer system 1100 can send and receive messages, data, information, and instructions, including one or more programs (i.e., application code), via the communication link 1122 and the communication interface 1112. The received program code may be executed by the computer processor 1104 once it has been received and / or stored in the disk drive component 1110 or any other non-volatile storage component for execution. The communication link 1122 and / or the communication interface 1112 can be used for electronic communication, for example, with a lesion assessment system 100, a computing platform 102, data storage 104, a display system 106, an image processor 108, and so on.

[0109] Where applicable, the various embodiments provided herein can be implemented using hardware, software, or a combination of hardware and software. Where applicable, the various hardware and / or software components described herein can be combined into composite components comprising software, hardware, and / or both without departing from the spirit of this disclosure. Where applicable, the various hardware and / or software components described herein may be separated into subcomponents comprising software, hardware, or both without departing from the scope of this disclosure. Furthermore, where applicable, software components may be implemented as hardware components, and vice versa.

[0110] The software relating to this disclosure, such as computer program code and / or data, may be stored on one or more computer-readable media. It is also conceivable that the software identified herein may be implemented using one or more general-purpose or dedicated computers and / or networked and / or other computer systems. Where applicable, the order of the various steps described herein may be modified, combined into composite steps, and / or separated into substeps in order to provide the features described herein. It is understood that at least some of the lesion assessment system 100, computing platform 102, data storage 104, display system 106, image processor 108, etc., may be implemented as such software code.

[0111] While this instruction is described in relation to various embodiments, it is not intended to be limited to such embodiments. On the contrary, this instruction includes various substitutes, modifications, and equivalents, as will be understood by those skilled in the art.

[0112] In describing various embodiments, this specification may present methods and / or processes as a specific set of steps. However, unless a method or process relies on a specific sequence of steps described herein, the method or process should not be limited to the specific sequence of steps described herein, and those skilled in the art will readily understand that the order may vary and still remain within the spirit and scope of various embodiments.

[0113] VII. List of various embodiments of this disclosure Embodiment 1: A method comprising receiving a set of fundus autofluorescence (FAF) images of a retina having one or more geographic atrophy (GA) lesions; receiving one or both of a set of infrared (IR) images of the retina or a set of optical coherence tomography (OCT) images of the retina; and using the set of FAF images and one or both of the set of IR images or OCT images to generate a GA lesion segmentation mask containing one or more GA lesion segments corresponding to one or more GA lesions in the retina.

[0114] Embodiment 2: The method according to Embodiment 1, further comprising using a processor to extract features of one or more GA lesions in the retina from a GA lesion segmentation mask.

[0115] Embodiment 3: The method according to Embodiment 2, further comprising a processor generating recommendations for treating one or more GA lesions based on extracted features.

[0116] Embodiment 4: The method according to Embodiment 2 or 3, wherein the extracted features include one or more GA lesions.

[0117] Embodiment 5: The method according to any one of Embodiments 2 to 4, further comprising combining one or more GA lesion segments into a single lesion component by a processor, wherein the extracted features include one or more of the area, perimeter, ferret diameter, or excess rim intensity of the single lesion component.

[0118] Embodiment 6: The method according to any one of Embodiments 1 to 5, wherein generating comprises generating a GA lesion segmentation mask using a neural network, the neural network comprising a U-Net deep learning neural network having an encoder and a decoder.

[0119] Embodiment 7: The method according to Embodiment 6, further comprising generating an encoded image input by concatenating a set of FAF images with one or both of a set of IR images or a set of OCT images using an encoder.

[0120] Embodiment 8: The method according to Embodiment 7, wherein generating a GA lesion segmentation mask includes decoding an encoded image input using a U-Net deep learning neural network decoder.

[0121] Embodiment 9: The method according to any one of Embodiments 1 to 8, wherein generating a GA lesion segmentation mask using a neural network, the neural network comprising a Y-Net deep learning neural network having a first encoder, a second encoder or a third encoder (either or both), and a decoder.

[0122] Embodiment 10: The method according to Embodiment 9, further comprising generating an encoded FAF image input using a set of FAF images via a first encoder, generating an encoded IR image input from a set of IR images via a second encoder, or generating an encoded OCT image input from a set of OCT images via a third encoder, or both of these.

[0123] Embodiment 11: The method according to Embodiment 10, further comprising generating an encoded image input by concatenating an encoded FAF image input with one or both of an encoded IR image input or an encoded OCT image input, and generating a GA lesion segmentation mask, which comprises decoding the input encoded image using a decoder of a Y-Net deep learning neural network.

[0124] Embodiment 12: A system comprising non-temporary memory and a hardware processor coupled to the non-temporary memory and configured to read instructions from the non-temporary memory and cause the system to execute the methods of Embodiments 1 to 11.

[0125] Embodiment 13: A non-temporary computer-readable medium (CRM) having recorded program code, wherein the program code includes code for causing a system to perform the methods of Embodiments 1 to 11.

Claims

1. A method performed by a processor, Receiving a set of fundus autofluorescence (FAF) images of the retina having one or more geographic atrophy (GA) lesions, Receiving one or both of the set of infrared (IR) images of the retina or the set of optical coherence tomography (OCT) images of the retina, Using one or both of the set of FAF images and the set of IR images or the set of OCT images, a GA lesion segmentation mask is generated which includes one or more GA lesion segments corresponding to one or more GA lesions in the retina. Methods that include...

2. The method according to claim 1, further comprising extracting the features of one or more GA lesions in the retina from the GA lesion segmentation mask.

3. The method according to claim 2, further comprising generating recommendations for treating one or more GA lesions based on the extracted features.

4. The method according to claim 2, wherein the extracted features include the number of the one or more GA lesions.

5. The method according to claim 2, further comprising combining one or more GA lesion segments into a single lesion component, wherein the extracted features include one or more of the area, perimeter, ferret diameter, or excess rim strength of the single lesion component.

6. The method according to claim 1, wherein the generation includes generating the GA lesion segmentation mask using a neural network, the neural network includes a U-Net deep learning neural network having an encoder and a decoder.

7. The method according to claim 6, further comprising generating an encoded image input by concatenating the set of FAF images with one or both of the set of IR images or the set of OCT images using the encoder.

8. The method according to claim 7, wherein generating the GA lesion segmentation mask includes decoding the encoded image input using the decoder of the U-Net deep learning neural network.

9. The method according to claim 1, wherein the generation comprises generating the GA lesion segmentation mask using a neural network, the neural network comprising a Y-Net deep learning neural network having a first encoder, a second encoder or a third encoder, or both, and a decoder.

10. Using the set of FAF images via the first encoder, an encoded FAF image input is generated, To generate an encoded IR image input from the set of IR images via the second encoder, or To generate an encoded OCT image input from the set of OCT images via the third encoder. To do one or both of the above The method according to claim 9, further comprising:

11. The method further includes generating an encoded image input by concatenating the encoded FAF image input with one or both of the encoded IR image input or the encoded OCT image input, The method according to claim 10, wherein generating the GA lesion segmentation mask includes decoding the encoded image input using the decoder of the Y-Net deep learning neural network.

12. It is a system, Non-temporary memory and A system comprising: a hardware processor coupled to the non-temporary memory and configured to read instructions from the non-temporary memory and cause the system to execute the method according to any one of claims 1 to 11.

13. A computer program that, when executed on at least one processor, causes the at least one processor to perform the method according to any one of claims 1 to 11.

14. A non-temporary computer-readable medium (CRM) storing the computer program described in Claim 13.

Citation Information

Patent Citations

  • Systems and methods for the detection of eye diseases

    JP2019531817A

  • Method and system for analysing images of a retina

    US20210319556A1

  • A patient tuned ophthalmic imaging system with single exposure multi-type imaging, improved focusing, and improved angiography image sequence display

    WO2020188007A1

  • Detection, prediction, and classification for ocular disease

    WO2020210891A1